Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Introduction

yaffle logo

yaffle

Yet Another File-Format Language & Engine: a language for binary file formats, and a compiler that generates readers and writers for TypeScript, Rust, Python, Go, C#, Java and C++.

  • The language: the .yfl reference.
  • Architecture: how the compiler, runtimes and tooling fit together.
  • IR: the contract between the compiler and its backends.
  • Targets: the generated API of each target language.

The yaffle language

yaffle describes binary file formats in .yfl files. A description is read in both directions: it parses bytes into values and serializes values back into bytes, byte-exact when the description says enough. docs/IR.md is normative for the precise semantics, and docs/CONFORMANCE.md defines the canonical JSON form and the error codes.

1. Principles

  1. Read-side expressions see only earlier fields. Lengths, at, conditions and attribute arguments may refer to fields declared before them, parameters, enclosing structs (lexical nesting) and bound $values. This one rule also explains endianness switches and every other scoping question.
  2. Computed fields (= expr) see every field, because the whole value is known at write time.
  3. Bytes that would read back differently are never written. A contradiction between a formula and a read-side expression is a compile error when provable, otherwise a write error.
  4. Fields fill themselves in where they can. A length, count, offset or size that can be recovered from what it describes is derived on write and drops out of the input.
  5. Byte-exact rebuilds are supported, not required. The defaults produce valid files. layout, @ref, split points and codec reuse exist to reproduce originals exactly.

Every struct therefore has two value shapes. The output is everything a parse produces. The input is what a serialize needs. Each field is one of:

FieldExampleOn writeIn the input?
inputu32 versionwritten as givenyes
derivedu32 count + T entries[count]recovered from what it describesno
computedu32 crc = crc32(body)its formulano
constantu16 magic = 42the constantno
optionalu32 align ?= 0x100as given, else the defaultoptional

Padding and @hidden fields are not in the input either. Within one parse, every struct read at a given address with a given type is one object, so shared pointer targets are shared values and cycles are allowed. Trailing bytes after the root are ignored.

2. Basics

2.1 Structs, primitives and blocks

struct Entry {
  u16   handler
  u16   kind
  u32   align
  u64   offset
  u64   size
  u32be crc                 // fixed endianness
}
  • C style: type before name, array size after the name. Semicolons are optional, so a struct pasted from Ghidra mostly parses as is. Comments are // and /* */.
  • u8, u16, u24, u32, u40, u48, u56, u64 (and i…), f16/f32/f64, bool, with le/be suffixes for fixed endianness. bool is one byte: nonzero reads as true (strict: only 0 or 1), and writes 0/1.
  • Blocks { … } group fields under attributes or conditions. Their fields flatten into the parent.

2.2 Arrays

struct Table {
  u32   count                      // derived: entries.length
  u32   pairs                      // derived: ids.length / 2
  Entry entries[count]
  u16   ids[pairs * 2]
  u8    hash[0x20]                 // a byte string
  u8    rest[]                     // to the end of the enclosing region
  u16   list[until 0xffff]         // terminator: consumed, not in the value, written back
  Bank  banks[until it.next == 0]  // inclusive: the element that satisfies it is part of the array
  str   pool[before ""]            // stops before an element equal to "", without consuming it
}
  • Derived lengths: a field used as a length, directly or through + − × << with constants, is derived from the array on write and left out of the input. Divisibility is checked.
  • Not derivable ((n + 7) / 8, rows * cols): the field stays an input and is checked against the array. A computed field (§4.1) can supply the formula instead.
  • Shared counts: several arrays using the same count must agree on write.
  • T x[] with struct elements repeats until the region ends. With an at over $index, each element goes at its own offset (§6.4).
  • until takes a terminator value, or a boolean expression over it (the element just read).

2.3 Strings

char   name[0x20]   // fixed: text up to the first \0, padded with \0 on write
str    path         // char[until 0]
str16  wide         // char16[until 0], UTF-16, follows $endian
str32  wider        // char32[until 0], UTF-32, follows $endian
u8     len
char   title[len]   // lengths count code units, never characters

char is UTF-8 by default; @encoding("latin1" | "ascii") changes it. Other encodings go through extern. Data after the \0 in a fixed char[n] is dropped by a parse (views keep it); use u8[n] if it matters.

2.4 Constants and checks

u16  magic = 42                   // constant: written automatically, checked on read
char tag[4] = "ARCV"
char order[2] in ("II", "MM")     // input type "II" | "MM"
u32  version in (1, 2, 5..9)
str  ver in /v\d\.\d{1,2}/        // portable regex subset (RE2-style: no backrefs or lookaround)
u32  size where it % 4 == 0       // any condition; `it` = this field's value

Checks apply in both directions: a violation is a CHECK error on parse and on serialize.

2.5 Literals

  • numbers: 0x4c43_4150, 0b1010, 1_000;
  • hex bytes x"5f2a 91c3" and base64 bytes b64"…";
  • four-character codes 'ARCV': the bytes in written order, whatever the endianness;
  • strings "…".

3. Attributes

3.1 How attributes work

An attribute is @name or @name(args). Where it goes depends on what it applies to:

struct Header @padUntil(0x80) {   // a struct: between the name and `{`
  u32 a @align(4)                 // a field: after the declaration
  @endian(be) {                   // a block: before the `{`
    u32 b
  }
  u64 c @align(8)
        @padAfter(8)              // a line starting with @ continues the previous field
}

When several apply, the one closest to the thing applies first, and each wraps the result so far: on a field that’s the leftmost, on a block the one just before {. So T x @size(exp) @via(zstd) @size(comp) is the decoded size, then the codec, then the size on disk. Arguments are read-side expressions (principle 1), and named arguments may also name positional ones (@align(to: 8)).

3.2 Endianness

struct TiffHeader {
  char order[2] in ("II", "MM")
  @endian(order == "MM" ? be : le) {
    u16 magic = 42
    u32 firstIfd
  }
}

There’s no file-wide setting. A struct inherits endianness from where it’s used, and the root defaults to le. @endian(x) sets the built-in bound value $endian (§5.2) for a struct, field or block; a le/be suffix on a type fixes it for one field.

3.3 Padding and alignment

struct Entry {
  u32 a        @padBefore(4)
  u32 b        @padAfter(4, 0xcc)        // fill byte
  u8  data[16] @align(0x100)             // pad before, from the nearest @base
  u8  tail[]   @align(0x10, i => i & 0xff) @strict
  str name     @padUntil(0x20, 0xff)     // "hello\0" then 0xff up to 0x20 bytes
  u8  head[25] @align(4, offset: 0x19)   // pad until (position + 0x19) % 4 == 0
  u8  blk[64]  @align(0x100, from: file) // from the file start, not the nearest @base
}
  • @align(n, fill?, strict?, from: file | base, offset: k).
  • @padUntil(n, fill?) pads after the content; content that is too large is a CHECK error on parse and an INPUT error on serialize. @padUntil(n, strict) is zero fill, checked.
  • @alignEnd(n) pads the end of a struct; on a @base struct, after all its targets.
  • Reading skips padding without checking it; @strict makes a fill mismatch a parse error.

3.4 Sizes and codecs

extern codec zstd
extern codec inflate @streaming                       // finds its own end
extern async codec oodle                              // the host's implementation is async
extern codec chacha20(u8 key[32], u64 nonce) @ranged  // can decode any slice, given its position

struct Packed {
  u32  expSize
  u32  compSize
  Asset body @size(expSize) @via(oodle) @size(compSize)    // decoded size · codec · size on disk
}

struct Node {
  u64     nonce
  u32     size
  Archive file @via(zstd) @via(chacha20(KEY_A, nonce)) @size(size)
}
  • @size measures whatever it wraps, so the inner and outer sizes are the decoded and on-disk sizes, both derived. The inner size is handed to the codec (Oodle and LZ4 need it). An @size on the wrong side of a non-streaming codec is a compile error.
  • async codec makes every schema containing it async.
  • @ranged: lazy views decode only what they read. Other codecs decode the whole region once.
  • Stable re-encoding: a codec is re-run only if its decoded content changed.
  • Partial codecs (e.g. one that unmasks only [0, 0x800)) take the range as parameters of a struct-level @via.

3.5 Other attributes

AttributeOnMeaningSee
@strictfield, block, struct, bitsfills, reserved bits and bool are checked on read3.3
@hiddenfieldnot in the input or output; derived or computed6.4
@encoding("latin1" | "ascii")char fields, aliasestext encoding2.3
@openenumunknown values pass through as numbers4.2
@bitorder(msb)bits, structbit order within the storage unit5.3
@basestructoffsets inside it count from its start6.1
@relativepointerthe offset counts from the pointer itself6.1
@nullable, @nullable(empty)pointer, at array (empty)a null value; an empty array as null6.1
@refpointerpoints at an equal value placed elsewhere6.3
@origin(expr)pointerthe offset counts from base + origin6.1
@split(fit | sizes)from fieldhow a stream is cut into pieces on write6.4

4. Values

4.1 Computed fields and externs

extern fn pathHash(char path[]) -> u32

struct Node {
  u32  n     = flags.length * 8     // the formula you choose where n isn't derivable
  u8   flags[(n + 7) / 8]
  u32  size  = sizeof(body)
  u32  crc   = crc32(body)
  u32  hash  = pathHash(path)
  u32  align ?= 0x100               // optional input: the default when absent, kept when given
  char path[0x40]
  Body body
}
  • = gives the raw value written, always computed. Without it the field leaves the input; with it it normalizes the input on write, and the field stays an input (§4.2). Formulas are not checked on read, except constants.
  • A formula that never agrees with the read side is an error, and one that agrees only for some values is a warning. A formula equal to what the field would be derived as is redundant. A consistent formula can still be lossy: a file with n = 5 above has one byte of flags and is rebuilt with n = 8, so it doesn’t rebuild byte-exact (§8).
  • Externs are declared in yaffle and implemented once per target language. Their failures are CODEC errors.

4.2 Transforms and enums

enum MemKind : u16 { Sys = 0, Vram = 1 }    // @open: unknown values pass through as numbers

struct Entry {
  MemKind kind
  u16     angle as it * 360.0 / 65536              // the inverse is worked out
  u8      name  as names[it]                       // inverse: names.indexOf(it)
  u32     tag   as decodeTag(it) = encodeTag(it)   // explicit inverse
  u32     hash  as hex(it)                         // one-way: input stays u32, read-only in views
  u32     flags = it | 0x80                        // normalize on write
}

as reads (raw → value, it = raw). = writes (value → raw).

4.3 Virtual fields

struct Node {
  u8  sizeHigh
  u32 sizeLow
  virtual u40 size = (sizeHigh << 32) | sizeLow    // no bytes; a write derives sizeHigh and sizeLow
}
  • A virtual field occupies no bytes. Its formula is evaluated on parse and inverted on serialize, so the parts it combines are derived and leave the input.
  • Invertible forms: bit concatenation (a << k) | b (where b fits in k bits), a * K + b (with 0 ≤ b < K), and linear forms, nested. Anything else is the error virtual-not-invertible.
  • Virtual fields take where/in checks and ?= defaults, but no attributes, at, from or as. A part may be a bit field, and an in (lo..hi) check narrows its range.

4.4 Derived arrays

struct Asset {
  u8  index[s] = indicesWhere(records, r => r.model != null)
  Key keys[n] where it == sortedBy(it, k => k.name)      // or a check stating the rule
}

Lambdas (x => expr) appear only as arguments of map, filter, indicesWhere and sortedBy. unique and concat take arrays. Lambda parameters may shadow fields.

5. Structure

5.1 Conditionals and unions

if (version >= 3) { u32 extra } else { u16 legacy }

switch (tag) {
  case "ARCV": Archive body
  case "ASST": Asset   body
  default:     u8      body[size]
}

union Node { Archive  Index }        // try in order; the variant's checks decide
  • A switch field’s value is a union discriminated by the tag. A union’s value records which variant matched; variants are named after their types.
  • if without else makes the block’s fields optional.

5.2 Parameters, nesting and bound values

struct Blob(u32 n) { u8 data[n] }          // explicit parameters, for reusable structs
struct Body(Header hdr) { … }              // any parameter type, structs included

struct Uses {
  Blob(16)    a                            // positional
  Blob(n: 32) b                            // named
}

struct Asset {
  u32    $version                          // bound: later fields and everything nested in them can read it
  Record records[r]
  struct Inner { … }                       // nested structs see outer fields lexically
}

struct MotionTable {                       // any depth below; the structs in between stay untouched
  if ($version >= 144) { u32 extra }
}

struct Node(u8 $key[32]) { … }             // bound parameters
struct Strict(u32 $version) { … }          // optional: pins the type, so the struct is checked on its own
  • $name is always a bound or built-in value. Plain names are only fields and parameters.
  • Nearest binding wins; inner bindings shadow outer ones (with a warning).
  • Every path is checked: a path that reaches a $version read without a binding is a compile error naming the path.
  • Built-ins: $endian, $offset (current position), $end (end of the enclosing region), $index (index in the enclosing array). A length using $end stays an input on write.
  • A root struct’s parameters, and the bound values it reads, are arguments to parse.

5.3 Types, generics and bitfields

type Fixed16 = i16 as it / 256.0
type Fixed(u8 bits) = i32 as it * 1.0 / (1 << bits)   // `* 1.0` makes it float division
type EncRec  = Record @via(xor(0x5a))        // a codec per element: `EncRec recs[n]`
type Name    = char[0x20]

struct Span<T> { u64 n  u64 ofs  T items[n] at ofs }   // generics: Span<Point>

extern type VarInt(u8 maxBytes) : u64        // a host-implemented primitive
extern type Half : f32 @size(2)              // fixed size

bits u8 {                       // explicit storage unit, LSB first; @bitorder(msb) flips it
  bool    compressed : 1
  MemKind kind       : 3
  u8                 : 4        // reserved: 0 on write, checked with @strict
}
u16 sizeHigh : 12               // C style: packs while consecutive fields share an integer type
u16 level    : 4
  • An extern type’s implementation provides size(bytes, ...args) (returning “need more bytes” when it can’t tell yet), read(bytes, ...args) and write(value, ...args). LEB128 (uleb128, sleb128) and common checksums are built in.
  • A bitfield’s storage unit follows $endian.

5.4 Modules and externs

import { Span, Name, SECTOR } from "./common.yfl"
import * as textures from "./textures.yfl"

export const u32 SECTOR = 0x800
export extern codec zstd                       // externs can be shared and imported by name
export struct Archive { u32 tag = 'ARCV'  … }  // a root: gets parse/serialize/view
  • export marks roots. Every top-level declaration can be imported, and exported structs of imported modules are roots too. Exported generic structs and aliases are importable but never roots. Import paths are relative, and .yfl may be omitted.
  • Cross-file cycles between types are allowed.

6. Placement

6.1 Pointers and bases

struct Asset @base {                      // offsets inside count from the start of Asset
  Model  *model : u64                     // the pointer is transparent: the value is just `model`
  Graph  *graph : i32 @relative           // target = address of this field + value
  Extra  *extra : u64 @nullable           // 0 → null; @nullable(0xffffffff) for other nulls
  Record *recs[count] : u64               // an array of pointers
  CurvePoint[n] *points : u64             // a pointer to an array; n is derived from its length
  Step[k] *steps : u64 @nullable(empty)   // an empty array writes 0, and 0 reads as empty
  u16[lens[$index]] *lists[9] : u64       // element-wise: lens[i] = lists[i].length on write
  str    *name : i32 @ref @origin(offsetof(pool))   // target = base + origin + value
  u64    tableOfs                         // derived from where `table` is placed
  Record table[count] at tableOfs
  u8     data[size] at sector * 0x800     // linear `at` → placement constraint (0x800-aligned)
}
  • Shared targets: the same offset gives the same object on read, and it’s written once.
  • at offsets that aren’t derivable stay inputs: the target goes exactly there, and overlaps are OVERLAP errors. Alternatively, compute it: = offsetof(t) - base.
  • An empty at array is placed at the current position, whatever its offset expression.
  • x == null / x != null test @nullable pointers.
  • Following a pointer is lazy in views, and an await point with async sources.

6.2 Layout: where pointer targets go on write

The default is depth-first declaration order: each struct comes before its targets. layout reproduces formats written in a different order:

struct Asset @base @alignEnd(0x10) {
  @padUntil(0x80) { /* header */ }
  str     pool[until ""]
  Record  records[r]        at recordsOfs
  u8      index[s]          at indexOfs
  Record *recPtrs[r] : u64  at recPtrsOfs @ref
  layout {
    records[*].model @align(0x80)        // pulled out: all model sets together
    records[*].anims @align(0x80)
    records[*]       @align(8)           // each record's remaining targets, in Record's own order
    records          @align(8)
    recPtrs
    index
  }
}
struct Record  { …  layout { curves, motions } }
struct Curves  { …  layout { x.points, y.points, x, y, this } }
struct Entry   { u32 align  u32 n  u64 ofs  u8 data[n] at ofs  layout { data @align(align) } }
struct World   { Volume volumes[n]  layout { volumes[*] { pieces[*]  portals[*].points } } }
  • layout orders only a struct’s own targets, using local paths. this is the struct itself. Targets go after the inline fields, and targets no item names follow the items in default order.
  • Pulling out: a container can pull shared groups out of inner structs (records[*].model), which overrides the inner layouts.
  • Item attributes may read fields, like computed fields. An item’s alignment applies to everything it places.
  • Groups (volumes[*] { … }) place, for each element, the named targets in order. The element itself comes first unless the group names this.
  • Fixed placements (at offsets that are inputs) go exactly where their offset says, and layout continues after them.

6.3 Shared values: @ref

struct Motion { str *name : u64 @ref }   // points at an equal string that's already placed

@ref pointers don’t own their target. On write they resolve to the input object itself if it was placed, otherwise to the first placed value with an equal encoding, searching the current base, then the enclosing ones. If none exists, the write fails with REF.

6.4 Streams

struct Fragmented {
  FragHeader hdr
  Frag       frags[] at max($index * 0x10000, 0x20)   // fragment k in its 64 KiB disk block
  Archive    file from frags[*].data @split(fit)      // joined decoded pieces, parsed as one stream
}
struct Frag {
  u32 magic = 0x46524147
  u32 exp                                             // decoded size: where the split falls
  u32 comp
  u32 check
  u8  data[] @size(exp) @via(zstd) @size(comp)
}

struct Overlay {
  u8     raw[0x40] @hidden
  Header hdr  from raw[0x00..0x20]                    // slices: raw is assembled from them on write
  Footer foot from raw[0x20..0x40]
}
  • from <stream> takes no space where it’s declared. A stream is a byte field, arr[*].field, or a slice raw[a..b]. Decoded bytes the struct doesn’t read are dropped.
  • A field consumed by from is derived on write and leaves the input; it stays in the output unless @hidden. @hidden fields must be derived or computed.
  • Splitting on write: the pieces’ size fields, when given in the input, win (byte-exact); otherwise the @split policy decides. sizes (the default) requires them. fit means the largest piece whose encoded form fits its block, and overlap checks enforce the fit.
  • Slice writes: the sliced field is assembled from every from slice of it, zero-filled elsewhere. Overlapping slices must agree (a warning when they provably overlap). It stays an input when it is also read directly, when its slices don’t provably cover it and it isn’t @hidden, or when its length depends on itself (u32 n u8 raw[n]: use [] with @size).

7. Built-ins

Built-inMeaning
sizeof(x), offsetof(x)size on disk, position (from the nearest base)
bytesof(x)bytes on disk (after codecs)
encode(x)plain encoding (before codecs)
crc32(x), adler32(x)checksums (of a field: its bytes on disk, like bytesof)
min, max, sumof values, or over one array (max(probes[*].total))
sizeof(this), offsetof(this)the current struct; for a @base struct, its whole region (write side)
x.lengthelement count; code units for strings
map, filter, indicesWhere, …derived arrays (§4.4)
$endian, $offset, $end, $indexsee §5.2
itthis field’s own value, in as, = and where
nullthe null pointer, in comparisons

8. Syntax details

  • Line rules. Members end at a line break, ;, or where the next member starts (u64 n u64 ofs). A line starting with @ continues the previous field unless it contains a { (then the attributes prefix a block). Lines starting with at, from, as, in, where or = also continue the previous field. Binary operators continue expressions across lines, and line breaks inside ()/[] don’t matter.
  • Operator precedence is Rust-like: * / % > + - > << >> > & > ^ > | > comparisons and in > && > || > ?:. So it & 0x80 == 0 means (it & 0x80) == 0.
  • Contextual keywords. type, from, at, … may be field names (Ghidra structs have type fields). Declaration names can’t be keywords, and field names can’t be it, true, false, le, be or this.
  • C conveniences: u8 r, g, b; struct Foo name; trailing };; Ghidra/C type names (uint32_t, dword, …) get a quick fix to the yaffle type. Multi-value cases: case 2, 3:. Enum members in expressions: Kind.Vram. Attributes may follow at (T x at ofs @size(n)).
  • Generic type arguments must be constants (Span<Blob(16)>, not Span<Blob(n)>), so an instance is context-free. Unions are inlined at each use.
  • Conditional fields can be read only where they’re known to exist: in the same branch, under a structurally identical guard, or when every branch declares them.
  • = with it (u32 flags = it | 0x80) normalizes the value on write, and the field stays an input. Without it it’s a computed field, or a constant (checked on read).
  • Integers are exact. Arithmetic has no overflow; values are range-checked where they reach a typed destination. Division truncates, and float-to-integer conversion rounds half away from zero.
  • Integer types from ranges. Integer as transforms and untyped constants get the smallest standard type holding the expression’s static range (u8 v as it + 100 → u16).
  • Formula analysis is by sampling the formula against the read side over the values the field can hold (its type, width and checks). Never agreeing is contradictory-formula, agreeing only sometimes is formula-mismatch, lossy is lossy-formula, and equal to the derived value is redundant-formula.
  • What makes a field derived: the first read-side use that is linear in the field (a length, an @size layer, an at offset, with placement constraints, or a struct argument: Blob(len) c derives len from Blob’s own derivation). Later uses are consistency checks.

9. Extensibility

LayeryaffleImplemented by
1. Transformsas expr, = expr, enums, type X = T as …yaffle expressions, extern fn
2. Codecs@via(codec), extern codec, async, @ranged, @streaminghost
3. Custom primitivesextern type X(args) : T with size/read/writehost

Architecture

.yfl files ──► parser ──► checker ──► IR ──► backends: TS, Rust, Python, Go, C#, Java, C++
                  │                            └─► generated code + a small runtime per target
                  └──► language service ──► language server ──► VS Code extension
  • Frontend (packages/compiler/src/frontend): a hand-written lexer, an error-tolerant recursive-descent parser, binder, checker and derivation. yaffle’s grammar is newline-sensitive (optional semicolons, @ lines continuing the previous field, < for both generics and comparisons, C declarators like T *name[n] : u64), which suits a hand-written parser with full control over recovery and ranges. See compiler.md.
  • IR (packages/compiler/src/ir, IR.md): a JSON-serializable, language-neutral program. Generics are monomorphized and aliases inlined. Bound $values and references to enclosing structs’ fields become implicit parameters passed at every use site, and every value the writer computes (derived, computed and constant fields) is an explicit derive expression. Every semantic question is answered here; backends only translate.
  • Backends (packages/compiler/src/backends/<target>): generate(program, options) returns source files. They emit imperative code, Kaitai-style, rather than schema objects for an interpreter. Each target has a small runtime (streams, views, commit, codecs) that the generated code imports or inlines.
  • Language service (packages/compiler/src/service): the editor-facing API over the frontend. The language server and the VS Code extension sit on top of it. See tooling.md.

Repository

PathWhat
packages/compilerfrontend, IR, language service, CLI and every backend
packages/runtime, packages/runtime-<lang>the runtime each target’s generated code uses
packages/lsp, packages/vscodelanguage server and VS Code extension
conformancethe shared test suite, its runner, and test externs per target

Targets

TargetRuntimeParse, serialize, JSONViews, surgical writesAsync
TypeScriptpackages/runtime✓✓sources, codecs, awaitable view chains
Rustpackages/runtime-rust✓✓async codecs run synchronously
Pythonpackages/runtime-python✓✓asyncio sources, codecs and views
Gopackages/runtime-go✓✓context-aware codecs, blocking reads
C#packages/runtime-csharp✓✓sources, codecs and views
Javapackages/runtime-java✓✓views on virtual threads
C++packages/runtime-cpp✓✓ (file-backed views)—

Every target generates the same entry points under its own conventions:

TargetParse / serializeJSONViews
TypeScriptX.parse(bytes), X.serialize(x), parseAsynctoJson / fromJsonX.view(src), $patches, $commit
RustX::parse(&bytes), x.serialize(), parse_withto_json / from_jsonX::view(bytes), set_…, commit()
PythonX.parse(data, **params), x.serialize(), parse_asyncto_json / from_jsonX.view(src), _patches, _commit
GoParseX(data, opts), SerializeX(v, opts), …ContextXToJSON / XFromJSONViewX(data, opts), setters, Commit
C#X.Parse(bytes, opts), X.Serialize(x, opts), ParseAsyncToJson / FromJsonX.View(…), sync and async
JavaX.parse(bytes, options), X.serialize(x, options), parseAsynctoJson / fromJsonX.view(bytes), patches(), commit(…)
C++X::parse(bytes), X::serialize(x)to_json / from_jsonX::view(bytes), X::view_file(path)

The details per target are in targets/<lang>.md.

Project config

yaffle.json lists the sources and, per target, the output directory and where extern implementations live. CLI flags override it, and the language server reads the same file.

{
  "sources": ["*.yfl"],
  "targets": {
    "ts": { "out": "gen/ts", "externs": { "zstd": "./externs/zstd.ts" } },
    "rust": { "out": "gen/rust", "externs": { "zstd": "./externs/zstd.rs#ZSTD" } }
  }
}
yaffle check schema.yfl                             # diagnostics with code frames
yaffle build --project yaffle.json --target ts
yaffle ir schema.yfl --text                         # the IR the backends see
yaffle fmt --check *.yfl

Testing

  • Conformance (CONFORMANCE.md): one suite of (.yfl, bytes) → JSON, JSON → bytes and round-trip cases, run against every target through a per-target adapter, in parse mode and again through lazy views.
  • Per target: native runtime tests (cargo, unittest, go test, xUnit, a Java runner, a C++ runner) plus vitest tests that compile .yfl, generate code, and build and run it.
  • Frontend and tooling: unit and golden-IR tests, language server tests at protocol level against both a fake and the real compiler, and grammar tests for the extension.

yaffle IR: semantics

The types live in packages/compiler/src/ir/ir.ts. This document says what they mean. It is the contract between the frontend (which produces IR) and the backends (which turn IR into code). Both sides must follow it.

1. What the frontend has already done

  • Monomorphized. Generic structs are instantiated per distinct type argument list, type aliases are inlined, constants are folded, and imports are resolved. Every struct referenced anywhere is in program.structs.
  • Lowered scoping. Bound values ($version) and lexical references to an enclosing struct’s fields become implicit parameters (origin: "bound" | "outer"). Each use site passes them explicitly in IrStructType.args. $endian is the implicit endian parameter (origin: "endian", value type { kind: "endian" }). A struct with usesEndian: false may ignore it.
    • Parameter order. IrStruct.params is always: endian (default le) first, then the declared parameters in declaration order (explicit, and declared $ parameters with origin bound), then implicit bound parameters sorted by name, then implicit outer parameters sorted by owner and field. Every struct has the endian parameter, even with usesEndian: false.
    • Implicit names. A lowered $version is the parameter version (boundName: "version"), suffixed with _bound if a declared parameter already has that name. An enclosing struct’s field count is the parameter count (origin outer), suffixed with _outer on a collision. A use site passes a field of the current struct if it binds the value before the use (or is the enclosing struct), and the current struct’s own implicit parameter otherwise.
  • Resolved derivation. Every value the writer computes (derived, computed and constant fields) has a derive expression, except the bytes of streams and slice carriers (§3.7).
  • Normalized expressions.
    • Constants are folded where the operands are literals, including enum and const references and sizeof(Type) of fixed-size types.
    • A character code assigned to or compared with an integer whose endianness is inherited lowers to cond($endian == be, <big-endian value>, <little-endian value>). With fixed endianness it is a literal. Such a field is still deriveKind: "constant", with a check it == <same expr>.
    • Arithmetic (+ - * / % << >> & | ^, unary - and ~) always has type bigint or float (f64), and both operands have exactly the result type: the frontend inserts convert on integer fields. Comparisons don’t convert when both sides already have the same type. An integer literal compared with an int field takes that field’s type (version >= 3 compares two u32).
    • String lengths are never len(<string>). Derived string lengths and s.length use encodedLength(s, unit, encoding), divided by unit for 2- and 4-byte units. len is used for arrays and bytes.
    • x == null is isNull(x) and x != null is !isNull(x). There is no null literal.
  • Inferred value types. The value type of an integer transform is the smallest standard integer type (u8/u16/u32/u64, i8/…/i64) holding the read expression’s static range (interval analysis); unbounded or beyond 64 bits gives i64 (u64 if non-negative). The read expression’s own type stays bigint. Untyped constants (const N = 4) get a type the same way.
  • Complete bit units. Every IrBits unit’s fields cover the whole storage unit: bits not declared become an implicit reserved field (name: null), so strict checks are well defined.
  • Checked: names, types, bindings on every path, layout paths, invertibility, and layout cycles. A program that reaches a backend is valid.

2. Reading

2.1 Context

A read happens in a context with:

buffer / sourcethe bytes (sync or async, random access)
posabsolute position of the next inline byte
baseabsolute position of the current base. The file start (0) at the root, changed by @base structs
endabsolute end of the current region: the source size at the root, narrowed by @size layers and by the extent of a codec’s decoded output
indexthe index of the element being read, inside arrays
strictwhether strict checks are on (root option, struct strict, or a strict layer)
paramsthe struct’s parameter values

$offset = pos - base, $end = end - base, $index = index. All offsets in expressions are relative to the base. Absolute positions never appear in values.

2.2 Members in order

Members are read in declaration order. Each field’s value becomes visible to later expressions under its name. Blocks and switches don’t introduce a scope: their fields are visible after them, as if declared inline. A field from a branch not taken is “absent”: reading it in an expression is a frontend error unless the expression is guarded. The frontend guarantees this.

  • Inline field: read its type at pos and advance pos by the bytes consumed.
  • placement: evaluate offset and read the type at base + offset without moving pos. The region’s end stays the same. With emptyIsNull (an array target), a stored offset of 0 reads as an empty array.
  • from: build the stream (§2.8) and read the type from it as a fresh region (base = 0, end = stream length). pos doesn’t move.
  • virtual: the field occupies no bytes. Evaluate virtual.read (type bigint, earlier fields only) and range-check the value against the field’s declared type, which is an integer prim (RANGE, like any typed location, §6). Nothing is read with the prim.
  • Checks: after reading the value, evaluate every check with it = the value. A false check is a CHECK error (CONFORMANCE.md §Errors). Strict-only checks run only in strict mode.

2.3 Types

TypeRead
primlittle/big endian per endian (inherit = the endian param). u24…i56 are N-byte integers; signed values are two’s complement and sign-extended. f16 is IEEE half.
boolone byte; nonzero → true (strict: anything other than 0 or 1 is a CHECK error).
stringsee the doc comment in ir.ts. Invalid UTF-8/16/32 is a CHECK error. ascii rejects bytes ≥ 0x80. latin1 maps bytes 1:1 to U+0000–U+00FF.
bytescount bytes, bytes until a terminator byte, or bytes up to end.
arrayelements per length. until: read an element, stop if it equals the terminator (compared as values). before: stop before an element that would equal the terminator; the terminator is not consumed. untilPredicate: read an element, stop if the predicate holds for it (it = the element); that element is included. rest: read elements while pos < end. If an element would cross end, that’s an EOF error. With elementPlacement, element i is read at base + elementPlacement($index = i), and with rest, reading stops when that offset is ≥ $end. Elsewhere, $index is the index in the nearest enclosing array.
structevaluate args in the caller’s context, then read the struct’s members. If the struct is base, set base = pos (or the placement address) for its contents. Struct-level layers wrap the struct (they apply inside any use-site layers).
enumread the storage prim. An unknown value is a CHECK error unless open; open unknowns are returned as plain integers.
pointerread storage. Equal to nullValue → null. Otherwise target address = base + offset (owned and ref alike), or fieldAddress + offset if relative, where the offset is signed; with origin, it is base + origin + offset, where origin is evaluated like a computed field (§3.2: it may use offsetof of any field of the struct, also later ones, and offsetof(this)). Read the target there without moving pos. The target may be an array (T[n] *p). With emptyIsNull (always with nullValue: "0"), the target is an array and a stored 0 reads as an empty array, so the value is never null. Identity: within one parse, the same target address and the same target type yield the same value object.
externcall host size(view, ...args) with a view starting at pos. undefined means “need more bytes”: give a longer view, and at end that’s an EOF error. Then read(bytes[0..size], ...args). With fixedSize, skip size().
uniontry each variant at pos. The first that reads without error (checks included) wins. If none does, that’s a UNION error listing each variant’s error.
transformread inner, then evaluate read with it = the inner value.
layeredapply the layers from outermost (last) to innermost (first) around reading inner (§2.4).

2.4 Layers on read

Processed from the outermost layer inwards:

LayerRead behaviour
align(to)skip forward until (pos - origin + offset) % to == 0. Strict: the skipped bytes must equal the fill.
padBefore(n) / padAfter(n)skip n bytes before or after the inner content.
padUntil(size)read the inner content, then skip until it occupies size bytes. Content larger than size is a CHECK error.
alignEnd(to)after the inner content, skip until (pos - origin + offset) % to == 0.
size(n)narrow end to pos + n, read the inner content inside it, then set pos to the window end. The leftover is padding (strict: must be zero).
via(codec, args)take the encoded bytes: the window [pos, end), unless the codec is streaming, in which case it reports how many bytes it consumed. Decode them and read the inner content from the decoded bytes as a fresh region (base = 0, end = decoded length). Then pos advances by the encoded length. If an inner size layer exists, its value is passed to the codec as outputSize.
endian(v)evaluate v (outside the layer) and use it as the endian argument for everything inside.
strictturn on strict mode inside.
  • Alignment origin. align and alignEnd measure from the current base by default. With from: "file" they measure from the start of the current byte space: the root source, a codec’s decoded region, or a stream, i.e. whatever $offset would count from if there were no @base in between. offset (default 0) aligns the position that many bytes ahead. The frontend emits from only as "file" and offset only when it isn’t 0.
  • Strict padding. A padding layer (align, alignEnd, padBefore, padAfter, padUntil) with strict: true checks its own fill on read even outside strict mode. It may appear on field, block and struct layers.
  • Endian layers rebind $endian: inside one, both endian: "inherit" and every { kind: "param", name: "endian" } expression, including the endian argument passed to nested structs, mean the layer’s value.

Padding fill: { kind: "byte" }, or { kind: "fn" } evaluated per byte index (0-based within that padding run), keeping the low 8 bits.

2.5 Bits

Read the storage unit (prim, endianness per endian). Field value = (unit >> shift) & mask(width), sign-extended for signed prims, converted for bool/enum. Reserved fields (name: null) must be zero in strict mode.

2.6 Switch and blocks

  • Block: evaluate condition. If true, read members, otherwise else (if any). Apply layers around the members.
  • Switch: evaluate the discriminant, then read the members of the first case with an equal literal, else default. If nothing matches and there is no default, that’s a CHECK error.

2.7 Errors

Every error carries a path (field names and array indices from the root) and an absolute offset. Codes are in CONFORMANCE.md.

2.8 Streams

  • field: that bytes field’s value (the decoded bytes).
  • each: the bytes field of every element of the array field, joined. from.split is present only for each streams.
  • slice: [start, end) of another stream.

Lazy implementations may decode pieces on demand. If every piece has a fixed decoded size, or the sizes are known from fields, an offset maps directly to a piece.

3. Writing

3.1 Input values

The writer takes an input value shaped like the generated input type: every field with input: "required", optionally fields with input: "optional", and never fields with input: "absent". If an absent field is present in the value, it is ignored. That way serialize(parse(bytes)) works directly.

The field kinds of DESIGN.md §1 map to these properties: an input field is input: "required" (with a derive only when it is normalized); a derived field has deriveKind: "auto" (stream split points are also input: "optional"), except that fields consumed by a stream and slice carriers are input: "absent" without a derive (§3.7); a computed field has deriveKind: "formula" and input: "absent"; a constant has deriveKind: "constant" and input: "absent"; an optional field has its default as derive and input: "optional".

  • A derive that uses it normalizes the input (u32 flags = it | 0x80): it is the field’s input value, the field stays input: "required", and the writer stores the value of derive. A derive without it comes with input: "absent" or "optional".
  • A default (u32 align ?= 0x100) is input: "optional" with the default as derive, deriveKind: "formula" and no check.
  • A virtual field’s value comes from the input (required, or optional with its default as derive). It is inverted on write, so its parts are derived from it: their derive expressions use >>, &, /, %, + and - on the virtual field’s value, with deriveKind: "auto" and input: "absent" (parts may be bit fields). A virtual field is never a part of another one. After writing, virtual.read must give back the input value, otherwise it’s a DERIVE error.

3.2 The write context and derive

derive expressions are evaluated in the write context:

  • field references read the input value, or the value of derive for fields that have one. Such fields may reference each other as long as there is no cycle (the frontend checks this).
  • sizeof / encodedSize / bytesof / encode take the field’s final encoding.
  • offsetof takes the field’s final position (from the base).
  • $index is the element index, and params are as on read.

A derived field’s final value can depend on positions that are only known after layout, e.g. u64 tableOfs derived from offsetof(table). Recommended writer shape (relocation style):

  1. Encode bottom-up into nodes: inline bytes plus a list of fixups (position, width, endianness, closure computing the value from final sizes and positions) plus a list of placeable targets.
  2. Run layout (§4) to assign every target a position.
  3. Evaluate and patch the fixups.
  4. Concatenate.

The frontend rejects layout cycles, i.e. a derived value whose own size depends on a position. Variable-size derived fields (varints, extern types) that feed into positions are a compile error.

3.3 Consistency

After writing, every field with a derive must satisfy its checks. The read-side expressions must agree with what was written: an array written with N elements whose length expression evaluates to M ≠ N is a DERIVE error, and so is a constant that would read back differently. A transform with write must round-trip: read(write(v)) == v, or it’s an INPUT error. Backends may skip the round-trip check for float transforms when the difference is below the half-ulp of the storage.

An untilPredicate array must end with an element satisfying the predicate, and no earlier element may satisfy it. A before terminator is not written.

3.4 Pointers on write

  • owned: the target is a placeable in the current base (the nearest enclosing @base struct, or the root). The pointer value is computed in a fixup: offset = targetPos - base (relative: targetPos - fieldPos; with origin: targetPos - base - origin), and a null value writes nullValue. With emptyIsNull, an empty target array writes nullValue (0); the same holds for the offset of a placement with emptyIsNull.
  • Identity: two pointers to the same object (reference identity in the input) share one target. Deduplicating equal values happens only for ref pointers.
  • ref: the pointer gets the offset of the input object itself if it was placed, otherwise of an already-placed, equal value of the same type: an element of an inline array, a field, or another target. Equality is equality of encodings. The current base is searched first, then the enclosing ones. If no such value exists, that’s a REF write error. When several match, the first placed wins.

3.5 Layers on write

These mirror reading:

  • size: written as the derived value (the frontend has derived the size field from encodedSize). A size that doesn’t match the actual inner length is a DERIVE error, unless the inner content is shorter and the remainder is padding.
  • via: encode the inner bytes with the codec. With an inner size layer, pass outputSize.
  • padUntil: content larger than the size is an error.
  • Fill bytes come from fill.

3.6 Codec reuse (byte-exact rebuilds)

If a value came from a parse (same process, tracked by object identity) and the decoded content of a via region is unchanged, the writer reuses the original encoded bytes instead of re-encoding. Views get this naturally. The same applies to from pieces.

3.7 Streams on write

  • from field: the consumed bytes field is input: "absent" and derived from the field read from it. It stays an output field unless hidden.
  • from arr[*].piece: in the element struct of arr, piece is input: "absent" (its content comes from the stream), and fields whose derive measures it (encodedSize(piece, i)) are input: "optional": split points, used when given for byte-exact rebuilds, otherwise from.split decides.
  • Slice carriers (carrier: "slices"): a bytes field read only through from slices of itself (T x from raw[a..b]) is input: "absent". On write it is a zero-filled buffer of the field’s own length (or, for rest lengths, the largest slice end) with every slice field’s encoding written at its start. Overlapping slices must produce identical bytes, otherwise it’s a DERIVE error. Carriers are only single-level slices of a field stream, and their length is a constant, an input-only expression, or rest: the frontend never marks a carrier whose length or slice bounds depend on the carrier itself.

4. Layout: placing pointer and at targets

4.1 Placeables

Within a base, a placeable is the target of an owned pointer or a field with placement (unless fixed). Each placeable has a unit: its own inline bytes (“this”) plus the units of its own placeables, ordered by its struct’s layout (default: this first, then its placeables in declaration order, recursively, i.e. pre-order depth-first). A nested @base struct’s unit is self-contained: its placeables are laid out inside its own region.

A placement is fixed: false exactly when its offset expression only uses fields that the writer computes (derived or computed fields); constants, parameters and input fields make it fixed. constraint comes from derived linear offsets (at sector * 0x800 → modulus 2048, remainder 0).

4.2 Base region

[base inline bytes][placeables per the base struct's layout]. A base’s this is always first, because offsets count from it.

4.3 Layout items

Each placeable is placed at most once. The rules below are what the TS runtime’s layout.ts implements. The other runtimes port it, and the language service’s layout simulator (layoutSuggest.ts) mirrors it.

  1. Claims first. Before anything is placed, every placeable named by any item of a container’s layout is claimed by that item and removed from every other unit, even when an earlier x[*] item would otherwise reach it (pulling groups out).
  2. Then items in order:
    • A path naming a placeable field (curves, records[*].model) places the unit of each placeable it matches.
    • A path naming inline data or [*] elements (records[*] where elements are inline in the records target) places the remaining units of all placeables reachable from those elements, in their own layout order.
    • this places the struct’s own inline bytes.
    • A group (volumes[*] { … }, IrLayoutItem.group) runs per element of the array its path names (the path has no trailing each), in order. If the element is a placeable (a pointer or at element target), its own bytes come first unless a group item names this (as in a unit, rule 3); then the group’s items, relative to the element. Everything the nested items name is claimed by the group. Targets of elements the group doesn’t name are leftovers.
  3. this comes first in a unit unless an item names it.
  4. Leftovers (placeables no item placed) follow the items in default order.
  5. Empty targets (empty arrays and strings) are placed at the current position after alignment, and a writer accepts any stored offset for them.
  6. Fixed placements go exactly where their offset says. When layout reaches one, the running position moves to its end if that is further on.

4.4 Alignment and constraints

  • Item layers (align) apply to every structure placed by that item, unless a nested layout item specifies its own. Group layers apply to nested items without layers of their own.
  • Placement constraints (constraint) add an alignment-like requirement: the writer moves the position forward to the next value ≡ remainder mod modulus.
  • Gaps are filled with zero.

Layout items that read fields. Arguments of a layout item’s layers (@align(align), @align(1 << shift)) are evaluated in the write context of the struct that owns the layout: its fields (input values, or their derive), its parameters and $index where applicable, the same as its computed fields. They never see fields of the placed target, so recs[*].body @align(count * 2) uses the container’s count. Any alignment, constant or not, must be at least 1; 0 or a negative value is a RANGE error, even when the item places nothing (arguments are evaluated per instance, like computed fields). Arguments can’t depend on sizes or positions (sizeof, offsetof). Conformance: conformance/cases/layout-fields.

4.5 Fixed placements and overlaps

A fixed placement goes exactly at the given offset, and overlapping another placed byte range is an OVERLAP error. The same applies to elementPlacement arrays: element i goes at its computed offset, and must not overlap element i+1.

5. Codecs, extern functions and extern types (host side)

Host implementations are found through yaffle.json targets.<lang>.externs[name] (a module path per target). Shapes in TS:

// extern fn pathHash(char path[]) -> u32
export function pathHash(path: string): number;

// extern codec zstd / extern async codec oodle / … @ranged
export const zstd = {
  decode(input: Uint8Array, ctx: { args: unknown[]; outputSize?: number; position?: number }): Uint8Array, // or Promise for async
  encode(input: Uint8Array, ctx: { args: unknown[]; position?: number }): Uint8Array,
};
// @ranged: decode/encode may be called on any slice; ctx.position is the slice's offset within the encoded region.
// @streaming: decode returns { output: Uint8Array; consumed: number } and gets the rest of the region as input.

// extern type VarInt(u8 maxBytes) : u64
export const VarInt = {
  size(bytes: Uint8Array, maxBytes: number): number | undefined,
  read(bytes: Uint8Array, maxBytes: number): bigint,
  write(value: bigint, maxBytes: number): Uint8Array,
};

Built-in extern types: uleb128 (value u64) and sleb128 (value i64) have builtin: true, module: "" and exactly those ids. Every runtime implements them; they need no externs entry. Built-in codecs that need no host: none. Built-in functions (crc32, adler32, …) are implemented by each runtime.

6. Expressions: integers, floats and built-ins

  • Exact arithmetic. Integer expression arithmetic is mathematically exact: no wraparound. A result is checked when it lands in a typed location (array length, offset, field, argument). Out of range is a RANGE error.
  • Division truncates toward zero. % takes the sign of the dividend. Division by zero is a RANGE error.
  • Shifts: the shift count must be in 0..63. >> is arithmetic on negative values.
  • Bitwise operators (& | ^ ~) work on the 64-bit two’s complement representation.
  • Mixed int/float arithmetic converts to f64 (the checker inserts convert).
  • round rounds half away from zero.
  • Comparisons across int widths compare mathematical values.
  • Equality of composite values. == and != on arrays, strings and bytes compare by value, element-wise; bytes may be compared with an array of integers (element-wise, as numbers). On structs they compare canonical JSON equality.
  • TS representation: values of up to 32 bits (and u40/u48 when ≤ 2^53) are number; 64-bit fields are bigint. Generated code may compute in number when it can prove the range stays below 2^53, and must use bigint otherwise.

Built-ins beyond the obvious:

  • encodedLength(s, unit, encoding) returns bytes, so char16 t[n] derives n = encodedLength(t, 2, "utf16") / 2.
  • this is only the argument of sizeof (write context only; for a @base struct its whole region including targets and @alignEnd), offsetof (the struct’s start from the base), bytesof and encode. Its type is the current struct.
  • sum(arr) is the sum of an array of numbers (0 when empty), of type bigint (f64 for floats). min/max with one array argument give its smallest/largest element (the element type); an empty array is a RANGE error.
  • lengths(arr) is the element count of each element of an array of arrays (pointer targets are transparent), as an array: the derivation of u16[lens[$index]] *lists[9] : u64, where $index inside the element type is the index in lists.
  • isNull(x) takes one argument of a nullable type.
  • Derived arrays. map, filter, indicesWhere and sortedBy take an array and a { kind: "lambda" } as their second argument (the only place a lambda appears; its body sees the enclosing context). unique takes an array, concat one or more. Result types: map → an array of the lambda body’s type; filter, sortedBy, unique → the input’s type (bytes stay bytes); indicesWhere → an array of bigint; concat → bytes if every input is bytes, else an array of the unified element type. A derive or check operand of array type assigned to (or compared with) a bytes field converts element-wise, range-checked (RANGE). sortedBy keys compare as numbers, strings by code points, or bytes lexicographically, and the sort is stable.

7. Generated API (every backend)

For each root: parse, parseAsync (if async or for async sources), serialize, serializeAsync, safeParse, view (tier 2), toJson, fromJson. Names are adapted to each language’s conventions (Rust: parse, serialize, to_json, …). Canonical JSON is defined in CONFORMANCE.md. Root parameters (explicit + bound + endian) become parse/serialize options.

Output and input types

  • Output type: every field with output: true.
    • Fields of conditional blocks without else are optional.
    • if/else blocks and switches become discriminated unions where a discriminant field exists (TS: a union of object types). Otherwise the fields from both branches are optional.
    • Pointers have the target’s type, plus null if nullable.
    • Unions are { type: "<Variant>", value }.
  • Input type: the same, minus absent fields, with optional fields optional.

8. Worked examples

yaffle ir file.yfl prints the IR for real sources. These sketches show the shapes.

8.1 Derived count

struct Table { u32 count  Entry entries[count] }
{ "kind": "field", "name": "count", "type": { "kind": "prim", "prim": "u32", "endian": "inherit" },
  "derive": { "kind": "call", "callee": { "kind": "builtin", "name": "len" },
              "args": [{ "kind": "field", "name": "entries", "type": { "kind": "array", "element": { "kind": "struct", "struct": "Entry" } } }],
              "type": { "kind": "bigint" } },
  "deriveKind": "auto", "checks": [], "hidden": false, "output": true, "input": "absent" }
{ "kind": "field", "name": "entries",
  "type": { "kind": "array", "element": { "kind": "struct", "struct": "Entry", "args": [{ "kind": "param", "name": "endian", "type": { "kind": "endian" } }] },
            "length": { "kind": "count", "expr": { "kind": "field", "name": "count", "type": { "kind": "int", "bits": 32, "signed": false } } } },
  "checks": [], "hidden": false, "output": true, "input": "required" }

8.2 Endianness block

struct TiffHeader { char order[2] in ("II", "MM")  @endian(order == "MM" ? be : le) { u16 magic = 42  u32 firstIfd } }

This gives a block with layers [{ kind: "endian", value: cond(order == "MM", be, le) }]. magic has derive: 42, deriveKind: "constant", a check it == 42, and input: "absent". order has literals: ["II", "MM"], an in check, and input: "required".

8.3 Layered sizes

struct Packed { u32 expSize  u32 compSize  Asset body @size(expSize) @via(oodle) @size(compSize) }
  • body.type is layered(struct Asset, [size(field expSize), via(oodle), size(field compSize)]).
  • expSize.derive is encodedSize(body, 0) and compSize.derive is encodedSize(body, 2).
  • program.structs has Packed.async = true because oodle is async.

8.4 Bound value lowering

struct Asset { u32 $version  Record r }
struct Record { Motion m }
struct Motion { if ($version >= 144) { u32 extra } }
  • Motion.params gets { name: "version", origin: "bound", boundName: "version", type: u32 }, and Record.params gets the same, because Record passes it through.
  • In Asset, the field r has type { kind: "struct", struct: "Record", args: [endian, field version] }.
  • In Record, m passes param version.

8.5 at with a derived linear offset

u32 sector  u8 data[size] at sector * 0x800
  • sector.derive is offsetof(data) / 0x800.
  • data.placement is { offset: sector * 0x800, constraint: { modulus: "2048", remainder: "0" }, fixed: false }.

The yaffle compiler frontend

The frontend turns .yfl sources into the IR specified in IR.md, reports diagnostics, and powers the language service and the CLI. It lives in packages/compiler/src/:

DirectoryContents
frontend/lexer, parser, binder, checker and lowering, formatter, printers
ir/the IR types (ir.ts), traversal helpers (walk.ts), the validator
service/the language service behind service/api.ts
cli/the yaffle command (runCli in cli.ts, the executable in main.ts)
index.tsthe public API: compile, compileSource, loadProjectConfig, build, …

Pipeline

runFrontend (frontend/program.ts) drives one compilation:

  1. Load. Read the entry files and, following imports, every module they reach. Parse results are cached per path and text (ParseCache; the language service keeps one).
  2. Parse. lexer.ts produces tokens that record whether a line break precedes them (members end at line breaks) and collects comments separately. parser.ts is an error-tolerant recursive descent parser: missing pieces become undefined, empty identifiers or Error expressions, so broken files still give a tree.
  3. Bind. symbols.ts creates a symbol per declaration, parameter and field. A struct’s fields are flattened through blocks, branches, cases and bitfields; each records the conditional branches (frames) it sits in. Then imports are resolved and modules ordered (imports first).
  4. Check. The Elaborator (elaborate.ts) checks every declaration once. Struct bodies go through members.ts, expressions through expr.ts.
  5. Assemble. instances.ts builds the program from the export roots (see below).
  6. Report unused imports, then return the program (unless there are errors), the diagnostics and the elaborator, whose SemanticModel (model.ts) the language service reads.

Module paths in the IR are relative to the project root (rootDir) or else to the common directory of all modules. Paths given to the compiler are used verbatim (win32 drive letters, \ separators and the language server’s virtual paths keep their spelling); only . and .. are resolved (util.ts).

Checking and lowering

One code path both checks and lowers. A Ctx (ctx.ts) says where an expression is: the module and struct in scope, the side (read: earlier fields only; write: every field; const), the conditional frames and guards around it, what it means, and whether to report diagnostics and record language-service information.

  • Check mode runs once per declaration, reporting and recording; generic type parameters are opaque. A non-generic struct’s check result is its IR.
  • Lower mode lowers a generic instance with its type arguments substituted, silently. Type arguments must be constants, so an instance doesn’t depend on where it’s used.

members.ts lowers a struct body in two passes. Pass A lowers what the read side needs, in declaration order: types, arrays, pointers, transforms, layers, placements, streams, bitfields, blocks, if and switch. Pass B lowers what may reference any field: computed fields, defaults, explicit transform inverses, pointer origins and checks. Then:

  • virtual.ts inverts each virtual field, deriving its parts by flattening its formula into a mixed radix ((hi << 32) | lo, a * 1000 + b).
  • slices.ts decides which bytes fields are assembled from their from slices on write.
  • derive.ts derives fields from their read-side uses: array/string lengths, @size layers, at offsets (with placement constraints) and struct arguments, inverting expressions linear in one field (invert.ts). The formulas of computed fields are sampled against the read side to find contradictory, partial, lossy and redundant ones. It also decides which placements are fixed and reports layout cycles.
  • The layout block is lowered last, with its paths resolved through arrays and pointer targets.

Supporting modules: types.ts (primitives, value types, static sizes), range.ts (interval analysis for the value types of integer transforms and untyped constants), attributes.ts (the attribute registry used by the checker, completion and hover), diagnostics.ts (the registry of diagnostic codes and quick-fix helpers), printer.ts and irtext.ts (source and IR printers; irToText is yaffle ir --text and the golden-test format), and format.ts (the formatter, which keeps the author’s column alignment and is idempotent).

Program assembly

assembleProgram instantiates every struct reachable from the export roots (plus any extra roots a tool asks for), then runs the requirements analysis:

  • Each instance’s direct needs are the bound values ($version), enclosing structs’ fields and $index it reads. Needs propagate along uses to a fixpoint, unless the use site binds the value itself or sits inside an array.
  • A need that reaches an export root is a diagnostic with the full path (R → records (Record) → m (Motion)) and a quick fix that adds a root parameter.
  • Implicit parameters are appended in the order IR.md §1 specifies, implicit arguments are filled at every recorded use site, and the provisional names used during lowering (\0b:version, \0o:Owner.field) are renamed in place.
  • Finally it reports inline recursion, marks the pieces of from arr[*].piece streams, computes async and usesEndian, and collects the enums, externs and exported constants the structs use. Struct ids are qualified names (Asset.Record), module-prefixed only when two modules declare the same name; generic instances are Span<Point>.

IR utilities

ir/walk.ts visits and maps expressions, types, layers and members. ir/validate.ts (validateIr, exported from the package root) checks a program against IR.md: arguments match parameters, read-side expressions use only earlier fields, references resolve, bits fit, input and derive are consistent, and the program is plain JSON. The frontend’s tests run it on every program they produce.

Language service

createLanguageService(host) (service/index.ts) implements service/api.ts.

  • Analysis (analysis.ts): a Workspace holds open-file overlays and the project configuration and bumps a version on every change. An Analysis is one frontend run over the open files and the project’s sources, cached until the next change; it indexes the semantic model by declaration and position. Project sources are found by expanding the sources globs through ProjectHost.readDirectory when the host has it.
  • Locating (locate.ts): the path of AST nodes at an offset.
  • Features: hover (describe.ts: field types, offsets, sizes, byte order, layers, inverses, checks, bit diagrams, binding sites, keyword and attribute docs), completion and signature help (completion.ts: context from the AST refined by lexing the text before the cursor), navigation, references and rename (index.ts), and document symbols, semantic tokens, folding, inlay hints and code lenses (features.ts). Quick fixes come from the diagnostics.
  • Samples (sample.ts): decodeSample compiles the file with the chosen struct as an extra root, generates TypeScript with the TS backend into a temporary directory (cached per program, removed on dispose), imports it and walks a lazy view, annotating each field with its value, file offset and size. Generated modules import the runtime through a shim that re-exports the service’s own copy. getRootParameters describes the root’s parameters for prompting.
  • Layout suggestions (layoutSuggest.ts): suggestLayout decodes a sample, rebuilds the placement tree of each base region, explains the targets’ file order with layout items (shortest paths first), chooses alignments from the gaps, and minimizes the result against a simulation of the runtime’s layout engine (IR.md §4.3). It resolves to undefined when the schema’s own layouts already reproduce the sample.

CLI

yaffle check [files…] [--project yaffle.json]
yaffle build [--project yaffle.json] [--target ts] [--out dir] [files…]
yaffle ir <file> [--pretty | --text]
yaffle fmt [files…] [--check] [--stdout]

Without files, check and fmt use the project’s sources; --project defaults to the nearest yaffle.json above the working directory. Diagnostics are printed with the source line and a caret underline (colored on a TTY unless NO_COLOR or --no-color). Exit codes: 0 ok, 1 errors in the sources (or files to reformat with --check), 2 usage errors, 70 internal errors.

build compiles the project, runs the backends registered in backends/index.ts for each target, and writes their files. Extern module paths in yaffle.json are relative to the project and are passed to the backends relative to the output directory.

Tests

packages/compiler/test/{frontend,service,cli}, run with npx vitest run --project yaffle packages/compiler/test/frontend packages/compiler/test/service packages/compiler/test/cli.

FileCovers
lexer, parserevery token and construct, error recovery, spans
diagnosticsevery diagnostic code, positive and negative (the registry must be covered), quick fixes
designgolden IR (golden/design-*.txt) for the DESIGN.md examples, and the IR.md §8 examples
loweringgolden IR (golden/lowering-*.txt) for groups, alignment origins, virtual fields, derived arrays, slices
language, semanticsbound values, nesting, $index, transforms, placements, aliases, generics, roots, …
tiffthe TIFF fixture schema (golden and assertions)
corpusevery conformance schema compiles, validates and formats idempotently
validate, format, apithe IR validator, the formatter, the public API
service/*every language-service feature, workspace behaviour, sample decoding and layout suggestions, fuzzed edits
clievery command, in process and as a process

Known limits

  • Presence of conditional fields is decided structurally: a use is guarded only by the same condition text (in an if or ?:), or when every branch declares the field. There is no implication reasoning, and member access to conditional fields of other structs isn’t checked.
  • Formula consistency (contradictory, partial, lossy, redundant) is decided by sampling the measures 0…4096, not by proof.
  • Struct-argument derivations work one level deep for plain struct fields, not arrays of structs. Element-wise lengths support exactly lens[$index].
  • Virtual fields can’t themselves be derived from a later read-side use, and their parts must be fields of the same struct.
  • before and predicate until are only implemented for arrays (strings and bytes report not-implemented).
  • Layout paths don’t traverse unions. Layout item arguments can’t use sizes or positions.
  • Incremental analysis is per workspace: any change re-runs the semantic passes over all files (parse results are cached per file).
  • Hover offsets inside switches, unions and after alignment are symbolic approximations.
  • suggestLayout doesn’t infer layout groups (a sample only a group can explain is reported as inexpressible), assumes every instance of a struct follows one order, and models placement constraints only as explicit @align.
  • Samples annotate fields read from streams at their position in the stream, not in the file.

Editor tooling

yaffle ships a language server (packages/lsp, @yafflelang/lsp, executable yaffle-lsp) and a VS Code extension (packages/vscode) that bundles it. The language knowledge lives in the compiler’s LanguageService (packages/compiler/src/service/api.ts); the server is a thin, defensive protocol adapter over it, so any LSP client gets the same features.

Language server

Features

FeatureWhat it does
DiagnosticsPushed per file, debounced. Every open file of a project is checked, and with the project scope every file in the project’s sources too, so importers update when an import changes. A service may report problems in other files (at an import’s target, say); each diagnostic goes to the file it names. yaffle.json problems appear on the config file.
Quick fixesCode actions from the fixes the service attaches to diagnostics: did-you-mean renames, add or remove an import, remove a duplicate attribute or a formula, add [*], and more. When the client sends back a published diagnostic, the server recovers the service’s original object (with its fixes). Only yaffle’s own diagnostics reach the service, and context.only is honoured.
CompletionContext aware: types and keywords at a member start, attributes after @ (per target), $ bound values, members after ., layout paths, import lists, pointer storage types, clauses, top-level declarations, snippets. Snippet items become plain text for clients without snippet support.
HoverTypes, fields (also imported), enums and members, attributes, $ values, built-in functions, keywords. Also answers with the cursor just after a word. Markdown when the client renders it.
Signature helpBuilt-in functions, attributes and struct arguments, with the active parameter.
NavigationGo to definition, find references and document highlights (same-file references), across files.
RenameFields, structs and other declarations across files, with prepare-rename. Uses versioned documentChanges when the client supports them.
SymbolsDocument symbols (hierarchical, or flattened for older clients) and workspace symbols from every project.
Inlay hintsDerived values where an explicit = … would go (= entries.length), inferred transform inverses, slice assembly, decoded / on disk labels inside @size, and parameter names for positional struct arguments.
Code lensessize: 0x20 (32 bytes) above fixed-size structs, and needs: $version (u32) above structs that read bound values. The latter runs yaffle.showBindingSources, which peeks the fields and parameters that bind the value.
Semantic tokensFull, delta and range; the legend is the service API’s token types and modifiers.
FoldingBlocks, block comments, and runs of line comments or imports.
FormattingWhole-document formatting with the client’s tabSize / insertSpaces. Files with syntax errors are left alone.
On-type formattingTyping @ on a fresh line aligns it with the first postfix attribute of the field above (or its name, as the formatter does); Enter after an @ continuation line goes back to the field’s indentation. Pure text layout, done in the server.

Custom requests (types in packages/lsp/src/protocol.ts, also exported as @yafflelang/lsp/protocol):

RequestParamsResult
yaffle/showIr{ textDocument }The file’s IR as JSON (compiled with unsaved contents; 64-bit integers as strings) and its diagnostics.
yaffle/rootParameters{ textDocument, root }The parameters the root struct needs (explicit and $ bound), so a client can ask for args.
yaffle/decodeSample{ textDocument, root, sample, args? }Decodes a sample with the root struct and returns annotations: the field’s range, a short rendering of its value, and the offset and size in the sample.
yaffle/suggestLayout{ textDocument, root, sample, args? }A workspace edit (in changes form) with layout { … } rules that reproduce the sample’s target order and alignment, or null when the sample already follows the default layout.

sample is { uri } (a file the server reads) or { base64 }. args is canonical JSON (docs/CONFORMANCE.md). Malformed params are protocol errors; anything else that goes wrong (unknown root, a sample that doesn’t decode, no compiler) comes back in the result’s error string, a message for the user.

Behaviour

  • Sync. Incremental document sync, positionEncoding: utf-16 (the service’s encoding, so positions pass through unchanged). Every edit reaches the service immediately; only diagnostics are debounced.
  • Projects. A file belongs to the nearest yaffle.json above it, and each config gets its own service; files with none share an inferred project. Configs at workspace-folder roots are loaded at startup and follow workspace-folder changes. Unsaved editor contents are visible to every project. Non-file documents (untitled:) get virtual paths and join the first workspace project.
  • Watching. The server registers watchers for **/yaffle.json and **/*.yfl. Config changes reload the project; configs appearing or disappearing move open files between projects. On-disk changes go only to the projects whose service read or probed the file (or, for creations and deletions, whose directory contains it). Files a service read from outside the workspace get their own watchers when the client supports relative patterns.
  • Staleness and cancellation. Each request yields once so a pending cancellation or edit can arrive, then answers RequestCancelled if cancelled, or ContentModified if its document changed (position-based requests only). A diagnostics pass yields between files and stops as soon as a newer edit arrives.
  • Fault isolation. Every service call is guarded: an exception is logged (throttled) and answered with an empty result, malformed results are sanitized (VS Code rejects a whole response over one bad range), and a service that fails to start is reported once with window/showMessage. Nothing is sent after shutdown, and a closed connection never throws.

Architecture

ModuleRole
server.tsYaffleServer / startServer(connection, options): lifecycle and capabilities, document sync, projects and watching, diagnostics scheduling, the standard requests.
custom.tsThe yaffle/* requests.
projects.tsOne service per yaffle.json plus the inferred project; nearest-config lookup (cached); the host that overlays unsaved contents and records what each service read.
safe.tsSafeService (guarded service calls), ErrorReporter (throttled logging), loggers.
convert.tsService ↔ LSP types, with sanitizing, and snippet → plain text.
diagnostics.tsPublished diagnostics merged across sources (each project and each config check), re-published only when they change.
semanticTokens.tsLegend, validation, delta encoding and the per-document cache for deltas.
uri.tsURI ↔ path for both path flavours (testable on any OS), keeping the client’s URI spelling.
config.tsThe yaffle.* settings, and a fallback JSONC yaffle.json loader.
attributeAlign.tsOn-type alignment of @ continuation lines.
fs.tsInjectable filesystem (nodeFileSystem, MemoryFileSystem).
protocol.tsCustom requests, settings and command ids; dependency-free, shared with clients.
main.tsThe yaffle-lsp executable.

startServer takes the service factory and, optionally, the compiler’s compile (for Show IR), its loadProjectConfig, a filesystem, a path flavour, a logger and initial settings. main.ts wires in the compiler.

Running

yaffle-lsp [--stdio | --node-ipc | --socket=<port> | --pipe=<name>] [--service <module>]
node --conditions=development packages/lsp/src/main.ts --stdio     # from the sources

--stdio is the default. --service <module> (or YAFFLE_LSP_SERVICE) runs the server on another module exporting createLanguageService(host), and optionally compile and loadProjectConfig; the end-to-end tests use it to run on a fake service.

Settings

Pulled with workspace/configuration (section yaffle) at startup and on change; clients without it can push them with didChangeConfiguration or initializationOptions.settings.

SettingDefaultEffect
yaffle.diagnostics.delay250Milliseconds after the last edit before checking again (0 on open).
yaffle.diagnostics.scopeprojectproject: every file of the project; openFiles: only open files.
yaffle.inlayHints.enabledtrueInlay hints.
yaffle.codeLens.enabledtrueCode lenses.
yaffle.format.enabledtrueFormatting.
yaffle.alignAttributesOnType.enabledtrueOn-type alignment of @ lines.

VS Code extension

  • Language yaffle for .yfl files: comments, brackets, auto-closing pairs, a word pattern that keeps $bound, @attr and numbers with _ together, indentation and on-enter rules (doc comments, case:, @ continuation lines), and a JSON schema for yaffle.json.

  • Grammar source.yaffle, written in src/grammar.ts and generated into syntaxes/yaffle.tmLanguage.json by the build (a test checks the committed file is current). It covers every declaration, primitives with le/be, attributes and their arguments, $ values, all literals (hex and binary with _, floats, x"…", b64"…", FourCCs, escapes), regexes after in, lambdas, operators, [*] and layout blocks. Expressions (lengths, =, as, where, in, at, from, attribute arguments) are tokenized in their own regions, so rows * cols in a formula is not taken for a pointer field. Semantic tokens from the server refine it.

  • Snippets for structs (plain, exported, parameterized, generic), magic constants, enums, bits, unions, switch, if/else, endianness blocks, pointers, at, @via fields, imports, constants, type aliases, extern codecs, functions and types, and layout.

  • Commands (category yaffle):

    CommandWhat it does
    Show IROpens a read-only yaffle-ir: JSONC view of the active file’s IR beside it, refreshed on every .yfl save.
    Decode Sample File…Asks for a root struct (from the document’s symbols), a sample file and the root’s parameters (prefilled with the last answers; bytes as hex), then shows decoded values at the end of each field line, with a hover table of values, offsets and sizes. A status bar item clears them. Re-decodes on save; a failed re-decode keeps the values and flags them as stale.
    Suggest Layout from Sample…Same questions, then applies the suggested layout rules.
    Clear Sample AnnotationsRemoves the decoded values.
    Restart Language ServerRestarts the server (also done when a yaffle.server.* setting changes).
    Show Language Server OutputOpens the server’s log.
    yaffle.showBindingSourcesInternal: the needs: $… code lens’s peek.
  • Settings, besides the server’s above:

    SettingEffect
    yaffle.server.pathA server to run instead of the bundled one: a JavaScript module (run with Node over IPC; .ts with --conditions=development) or an executable (started with --stdio). Relative paths resolve against the first workspace folder.
    yaffle.server.runtimeThe Node executable for a module server; empty for VS Code’s own.
    yaffle.trace.serverTrace the LSP traffic in the output channel.
    yaffle.sample.redecodeOnSaveDecode the current sample again on every .yfl save (default on).

    The extension turns on editor.formatOnType and semantic highlighting for yaffle files.

Building and packaging

npm run build -w yaffle-vscode      # grammar JSON, dist/extension.cjs, dist/server.mjs
npm run package -w yaffle-vscode    # dist/yaffle.vsix

esbuild bundles the extension as CommonJS (what the extension host loads, vscode external) and the server with the compiler as ESM, with a createRequire banner because the bundled CommonJS dependencies require() Node built-ins. Both bundle straight from the TypeScript sources through the development condition, so no package needs a tsc build first. vsce package runs with --no-dependencies: everything is bundled, so runtime libraries are devDependencies.

Tests

npx vitest run --project @yafflelang/lsp --project yaffle-vscode
  • packages/lsp/test: the server over an in-memory JSON-RPC connection against a fake service (fakeService.ts, predictable text rules) with VS Code-like, minimal and plain-text clients and both path flavours; resilience against a throwing or malformed service; unit tests of the pure modules; main.ts over stdio.
  • packages/lsp/test/integration: the same server on the real compiler and a small invented project (project.ts): every feature at every word, keystroke-by-keystroke and random edits, workspace changes, Windows paths, sample decoding, and main.ts over stdio.
  • packages/vscode/test: the grammar through vscode-textmate (every construct and the DESIGN.md examples), the manifest, language configuration and snippets, building and vsce package, and the bundled extension activated against a vscode stand-in (vscodeMock.cjs) with the real language client and bundled server running every command.

Known gaps

  • Diagnostics are pushed; there are no pull diagnostics.
  • A project created to answer a request about a file that isn’t open (hover after go to definition, say) stays loaded until shutdown.
  • A yaffle.json created above the workspace folders is only noticed in directories a service has read from.
  • Imports from non-file (untitled:) documents don’t resolve.
  • No range formatting.
  • The grammar is heuristic where TextMate can’t know types: Type name decides a field declaration, and several fields on one line after an expression aren’t always split.
  • The extension isn’t tested in a real VS Code (that needs a download and a display); the activation test uses a vscode stand-in with the real client and server.
  • A missing executable in yaffle.server.path is reported by the language client itself; a missing module is checked up front.

Conformance suite

Every backend must pass the same cases: the suite is the oracle that keeps the seven targets identical. Expected JSON and bytes are computed independently of every backend, by hand or by the small generators in conformance/scripts/, never by running a backend.

1. Layout

conformance/
  cases/<area>/schema.yfl      entry schema (it may import other .yfl files in the directory)
  cases/<area>/cases.json      the cases (§2), generated by scripts/areas/<area>.ts
  cases/<area>/yaffle.json     optional: extern implementations per target (§5)
  externs/<target>/            the conformance externs for each target (§5)
  scripts/                     case generators, their byte/JSON library (lib.ts), gen-cases.ts
  src/                         runner, case loader, canonical-JSON comparison, reports, CLI
  test/conformance.test.ts     vitest entry: every case on every available backend
  test/*.test.ts               tests of the runner, the generators and the externs

Each backend provides an adapter at packages/compiler/src/backends/<target>/conformance.ts exporting adapter: ConformanceAdapter. The contract is conformance/src/adapter.ts:

export interface ConformanceAdapter {
  target: string; // "ts", "rust", …
  capabilities?: { views?: boolean }; // views: also run as "<target>:view" (§7)
  available(): Promise<boolean>; // is the toolchain usable here?
  run(input: AdapterRunInput): Promise<CaseResult[]>;
}
export interface AdapterRunInput {
  caseDir: string; // absolute
  program: IrProgram; // the case directory's schema.yfl, compiled by the frontend
  cases: Case[]; // normalized by the runner (below)
  workDir: string; // scratch directory for this target and case directory
  externs: Record<string, string>; // extern name → module for this target (§5)
  mode?: "parse" | "view"; // §7
}
export interface CaseResult {
  name: string;
  ok: boolean; // the adapter's own verdict
  message?: string;
  skipped?: string; // the target can't run this case (unsupported feature): reason
  json?: unknown; // actual canonical JSON (parse cases): always return it
  bytes?: string; // actual output hex (serialize, or the round trip of a parse case)
  error?: { phase; code; path?; offset?; message? }; // the error raised, if any
}

Adapters should be thin. They return json, bytes and error, and the runner re-checks them with its own comparison, so every backend gets the same diffs. bytes may be left out when it would be huge (large files); then only ok counts for the round trip. An adapter that throws an error whose message contains “not implemented” skips the whole case directory for its target; any other throw fails every case of the directory.

What adapters receive (Case, normalized by the runner): bytes is canonical hex (no whitespace) or { file: "<absolute path>" }; slices are already cut into a scratch file; expect and description are removed.

2. cases.json

{
  "description": "…",                   // optional
  "generatedBy": "conformance/scripts/areas/arrays.ts",   // optional
  "cases": [
    {
      "name": "basic",
      "root": "Table",                  // exported struct
      "args": { "version": 144 },       // root parameters (canonical JSON), optional
      "bytes": "0300000001020304",      // hex (spaces and _ allowed), or { "file": "input.bin" }
      "json": { "count": 3, "entries": [ … ] },   // expected parse result (canonical JSON)
      "roundtrip": true,                // default true (see below)
      "strict": false                   // parse in strict mode
    },
    {
      "name": "bad magic",
      "root": "Table",
      "bytes": "ffff",
      "error": { "phase": "parse", "code": "CHECK", "path": ["magic"] }
    },
    {
      "name": "write derives count",
      "root": "Table",
      "input": { "entries": [ … ] },    // serialize-only case: input JSON (count omitted)
      "bytes": "…"                      // expected output
    },
    {
      "name": "large sample",
      "root": "Archive",
      "bytes": { "file": "sample.bin", "offset": 768, "length": 4160320 },
      "expect": [ { "pointer": "/count", "equals": 6 }, { "pointer": "/names", "length": 598 } ]
    }
  ]
}
  • Parse cases (bytes + json and/or expect): toJson(parse(bytes)) must deep-equal json (§3), and every expect check must hold on it.
  • Round trips of parse cases:
    • true / "json" (the default): serialize(fromJson(json)) must equal bytes, where json is the actual canonical JSON (equal to the expected one when the case has it);
    • "value": serialize(parse(bytes)) must equal bytes in the same process, so codec reuse applies (IR.md §3.6). Used for codecs whose re-encoding differs from the original bytes;
    • false: no round trip (non-canonical input, e.g. non-zero padding or a NaN payload).
  • Serialize cases (input + bytes): serialize(fromJson(input)) must equal bytes.
  • Error cases: the phase and code must match, and the path too if given.
  • bytes: hex, or { "file", "offset"?, "length"? }. A file is relative to the case directory. offset/length cut a slice.
  • expect (runner-side checks on the actual JSON, for files too big to spell out): a list of { "pointer": "<JSON Pointer>", "equals"?: <canonical JSON>, "length"?: n, "byteLength"?: n, "startsWith"?: "<prefix>" }. length is an array’s or string’s length, byteLength a hex string’s byte count.

The case loader validates every cases.json (unknown keys, duplicate names, bad hex, missing files, error codes, mutually exclusive fields); an invalid directory fails all its cases. cases.json is never edited by hand: change the generator and run node conformance/scripts/gen-cases.ts.

Rules the cases pin beyond §3 and §4:

  • Bytes after the root struct are ignored by parse.
  • Computed fields are not checked on read (a wrong crc32 or size parses; the rebuild recomputes it), except constants.
  • A count larger than the remaining data is EOF, whatever its magnitude.
  • @ref resolves by identity first (the input object itself, e.g. restored by $ref), then by equal encoding (the first placed value wins).
  • On write, fields of an untaken conditional branch are ignored; a missing field of the taken branch is INPUT.

Known gaps:

  • The suite has no negative compile cases (diagnostics are covered by the frontend’s own tests) and no surgical-write or async-source cases (view mode covers reads and no-op commits).

3. Canonical JSON

This is the JSON form of values, used by toJson/fromJson in every backend and by the cases.

ValueJSON
structobject, fields in declaration order (fields from blocks, switches and bits inline), only output fields. Fields of untaken conditional branches are omitted. Derived, computed and constant fields are output fields.
integer ≤ 32 bits, or u40/u48/i40/i48number
56- and 64-bit integerdecimal string, e.g. "18446744073709551615", also when small ("1")
floatnumber. NaN, ±Infinity and -0 are the strings "NaN", "Infinity", "-Infinity" and "-0".
booltrue / false
bytes (u8 arrays only)lowercase hex string; other element types are JSON arrays
stringstring
arrayarray
enummember name; an open unknown value follows the integer rule of the storage (a number, or a decimal string for 56/64-bit storage)
pointerthe target’s JSON; null pointer → null
shared target (2nd and later occurrence, same object){ "$ref": "/records/1/model" }, a JSON Pointer (RFC 6901) to the first occurrence, in document order. Only for object- and array-valued targets (structs, unions, non-byte arrays); scalars, strings and bytes are always written out. A $ref may point at an ancestor (cycles).
union{ "type": "<Variant>", "value": … }
transformthe transformed value’s JSON (one-way transforms: output only; the input is the raw value)
extern typeJSON of its value type
bit fieldper its declared type (u64 lo : 40 is a decimal string)
bound field u32 $versionthe key without $ ("version")

Floats: all NaNs read as "NaN"; writing "NaN" gives the canonical quiet NaN (f16 7e00, f32 7fc00000, f64 7ff8000000000000). A number written to a narrower float rounds to nearest, ties to even. Key order is significant: the runner reports fields out of declaration order.

fromJson accepts this form. Fields that are absent from the input type are ignored if present, so fromJson(toJson(x)) works. $ref restores object identity; a dangling $ref is an INPUT error.

Root arguments (args): explicit parameters by name, bound parameters without $, and endian as "le" or "be" (default "le").

4. Errors

CodeMeaning
EOFnot enough bytes (including an extern type’s size asking for more at the end)
CHECKa check failed (constant, in, where, regex, strict padding or reserved bits, bool, enum, string encoding, switch without match), on read or on write
RANGEan integer out of range for its destination, division by zero, a negative length, an index out of range
UNIONno union variant matched
CODECa codec failed
INPUTinvalid input value on write: wrong JSON type, missing required field, string too long (or containing NUL), content larger than @padUntil, an element equal to an until terminator, unknown enum name or union variant, value missing from an indexOf inverse, a transform that doesn’t round-trip, a dangling $ref
DERIVEa derived value contradicts the data: shared counts disagree, an indivisible derived count, a non-derivable count or a constant length that doesn’t match the array, a size mismatch, piece sizes that don’t cover a stream
REFa @ref pointer has no equal placed value
OVERLAPfixed placements overlap (each other or inline bytes), or a stream piece doesn’t fit its block

Error objects expose code, path, offset (absolute, or -1 on write) and message. path lists field names and array indices from the root to the innermost field (or element) being read or written when the error happened. Blocks, switches and bits add no segment (their fields are flat); pointers are transparent. Examples: ["entries", 2, "x"], ["ifd", "count"]. A case leaves the path out where it is not determined (e.g. DERIVE between two fields, reserved bits, a switch without a match).

5. Externs in cases

A case directory whose schema declares externs has a yaffle.json mapping each extern to a host implementation per target, with paths relative to that yaffle.json:

{
  "sources": ["schema.yfl"],
  "targets": {
    "ts": { "out": "…", "externs": { "rle": "../../externs/ts/rle.ts" } },
    "rust": { "out": "…", "externs": { "rle": "../../externs/rust/rle.rs#RLE" } }
  }
}

The runner resolves them and passes the adapter its own target’s map (AdapterRunInput.externs): file paths (starting with ./, ../ or /) become absolute and keep an optional #ITEM suffix naming the item in that file; anything else (e.g. a Rust path like crate::hash::gt_hash) is passed through unchanged. A mapped file that doesn’t exist makes the directory invalid. The adapter makes generated code use them; each file in conformance/externs/<target>/ documents its target’s host shape.

TS: the module exports a value named exactly like the extern, in the shapes of IR.md §5 (export const rle = { decode, encode }, export function sum8(…), export const VarInt = { size, read, write }). Codec arguments arrive in ctx.args (u8 → number, u64 → bigint, u8[n] → Uint8Array).

Rust (the host interface of packages/runtime-rust): the file becomes a module of the generated crate and can use yaffle_runtime. The item is the extern’s name, or the #ITEM of the mapping. Dependencies other than the runtime are listed in yaffle.json under targets.rust.dependencies.

ExternItem
extern fn f(T a, …) -> Rpub fn f(a: T, …) -> R (ints as their Rust type, u8[] → &[u8], char[] → &str)
extern codec c(args)a value implementing yaffle_runtime::Codec: decode(&self, input, &CodecCtx { args, output_size, position }), encode(…); arguments are yaffle_runtime::Args
@streaming codecalso decode_streaming(&self, input, ctx) -> Result<(Vec<u8>, usize), String> (output, bytes consumed)
async codecthe same trait, synchronous: an async codec gives the same bytes as its sync twin
extern type X(args) : Ta value implementing yaffle_runtime::ExternType<Value = T>: size(&self, bytes, args) -> Option<usize>, read, write

conformance/externs/rust/stub/yaffle_runtime.rs mirrors that interface so the extern files can be unit-tested without the runtime (conformance/scripts/test-rust-externs.sh).

Errors returned by a codec are CODEC errors.

The conformance externs (conformance/externs/<target>/, one implementation per target):

ExternBehaviour
extern codec xor(u8 key) @rangedevery byte XOR key
extern codec xorpos(u8 key) @rangedbyte i XOR (key + i) & 0xff, i from the start of the region
extern async codec axor(u8 key)xor, asynchronous where the target has async codecs
extern codec rle(count 1..255, byte) pairs; greedy encoder; checks outputSize
extern codec rlez @streamingrle pairs then a 0x00 terminator
extern fn sum8(u8 data[]) -> u8sum of the bytes mod 256
extern fn swap16(u16 v) -> u16byte swap
extern type VarInt(u8 maxBytes) : u64unsigned LEB128
extern type Bcd16 : u16 @size(2)four BCD digits, most significant first

6. Running

all conformance testsnpx vitest run --project conformance
runner and generator testsnpx vitest run --project conformance test/runner.test.ts test/units.test.ts test/generated.test.ts
pass/fail matrixnode --conditions=development conformance/src/cli.ts [--filter …] [--targets …] [-v] [--json]
against other implementations… cli.ts --compiler <path/to/index.ts> --adapter <path/to/conformance.ts> (e.g. another worktree’s frontend or backend)
list / validate casesnode --conditions=development conformance/src/cli.ts --list / --validate
regenerate cases.jsonnode conformance/scripts/gen-cases.ts [<area>…] (--check to verify)
Rust extern unit testsnix shell nixpkgs#rustc --command conformance/scripts/test-rust-externs.sh
Environment variableEffect
YAFFLE_CONFORMANCE_FILTERcomma-separated patterns on <dir>/<case name>: substrings, globs with *, ! excludes
YAFFLE_CONFORMANCE_TARGETScomma-separated targets (ts,rust); ts also selects ts:view, ts:view only the view
YAFFLE_CONFORMANCE_VIEWS=0no view-mode targets (§7)

The CLI’s --filter and --targets set the first two.

7. View mode

Adapters that declare capabilities: { views: true } are also run as a second target named <target>:view (for example ts:view). In view mode the runner sends the same case directories with mode: "view" and skips serialize-only cases. For every other case, the adapter must:

  1. open the target’s lazy view over the case bytes (the bytes, or a random-access source over the file for { file } cases);
  2. produce json by reading every field through the view (materialize, then canonical JSON), not by calling parse;
  3. produce bytes by committing the view with no edits, which must reproduce the input exactly (the round-trip check), unless roundtrip is false;
  4. report errors raised while reading through the view (same code and path as parse mode).

The runner judges view results exactly like parse results. View mode exists because views have their own code paths (lazy offsets, decoded regions, streams) that parse-mode cases never exercise.

TypeScript target

The TS backend (packages/compiler/src/backends/ts) compiles the IR into one TypeScript module per .yfl module. Generated code imports the runtime @yafflelang/runtime (packages/runtime) as $rt. The output passes tsc --strict with exactOptionalPropertyTypes, noUncheckedIndexedAccess and erasable syntax only.

// yaffle.json
{
  "targets": {
    "ts": {
      "out": "src/gen",
      "externs": { "zstd": "./externs/zstd.ts", "pathHash": "./externs/hash.ts#hash" },
      "runtimeModule": "@yafflelang/runtime" // optional, the default
    }
  }
}

Generated API

Every exported struct X gets an entry-point object X:

import { Archive } from "./gen/archive.ts";

const pf = Archive.parse(bytes); // Archive
const out = Archive.serialize(pf); // byte-identical when nothing changed
const r = Archive.safeParse(bytes); // { ok: true, value } | { ok: false, error }
const p = await Archive.parseAsync(source); // any source, sync or async

const v = Archive.view(bytes); // lazy, editable
v.entries[3].align = 0x200; // same size: patched in place
v.entries.push(entry); // structural: re-encoded on commit
v.$patches(); // [{ at, remove, insert }] against the original bytes
const updated = v.$commit({ grow: "splice" });

const json = Archive.toJson(pf); // canonical JSON (docs/CONFORMANCE.md §3)
const input = Archive.fromJson(json); // ArchiveInput
MemberNotes
parse(source, options?)Uint8Array or sync source; reads the whole source.
safeParse(source, options?)YaffleErrors become { ok: false, error }; other exceptions propagate.
parseAsync / safeParseAsyncAny source.
serialize(value, options?)Takes an XInput (a parse result or a view works too).
serializeAsync(value, options?)Promise variant.
view(source, options?)Sync view for byte arrays and sync sources, async view for async sources.
viewAsync(source, options?)Always an async view.
toJson(value) / fromJson(json)Canonical JSON. Shared targets become JSON-Pointer $refs; one-way transforms show their raw value.
optionsFromJson(json)Options from canonical JSON (root parameters, endian, strict).
info{ name, async }.

Formats with async codecs (extern async codec) have no parse, safeParse or serialize, and view always returns an async view.

Options (XOptions extends $rt.BaseOptions): endian (the root $endian, default "le"), strict (padding fill, reserved bits and bool bytes are checked), plus the root’s parameters, bound parameters without their $. Parameters with a default are optional.

Errors are YaffleErrors with code, path (field names and indices from the root), offset (absolute, -1 on write), detail and, for UNION, causes:

CodeMeaning
EOFData ends before a read is complete, or a count can’t fit the region.
CHECKA constant, in/where check, enum member, strict padding or bool is wrong.
RANGEA value doesn’t fit its type, or a length, offset or alignment is invalid.
UNIONNo union variant matched.
CODECA codec or extern failed.
INPUTThe input value has the wrong shape or type, or misuse of the API.
DERIVEA given value disagrees with what the format derives (lengths, offsets, sizes).
REFA @ref pointer has no equal value placed in its base.
OVERLAPFixed placements overlap, or a stream piece doesn’t fit its block.

Types

Each struct gives an output type X and an input type XInput:

yaffleTypeScript
integers up to 48 bits, floatsnumber
56/64-bit integersbigint (bigint | number in XInput)
boolboolean
stringsstring
u8 x[…]Uint8Array
arraysT[]
enumsa union of member names (open enums add number)
pointersthe target’s type; T | null when @nullable
union{ type: "Variant"; value: T } | …
switch on a fielda discriminated union, narrowed on the tag’s literal values
if without else, branchesoptional fields
as transformsthe value type; XInput takes the raw type for one-way transforms
constants, in (…) literalsliteral types
extern typesthe declared value type

XInput leaves out derived, computed and constant fields, and makes ?= defaults and stream split points optional. Expressions compute in number where interval analysis proves they stay within ±2^53, otherwise in bigint; conversions to lengths and offsets are range-checked. User types named like a global the generated code uses (Record, Map, Error, …) get a _ suffix.

Internally each struct also gets read_X (and readAsync_X), scan_X (views), write_X (and writeAsync_X), toJson_X, fromJson_X and the view descriptor $view_X; these are internals shared with the runtime and may change.

Sources

A source is a Uint8Array, a SyncSource { size, read(offset, length), write?(offset, bytes) } or an AsyncSource whose read (and write) return Promises. A source is detected as async by a zero-length probe read. arraySource(bytes) and asyncArraySource(bytes) from @yafflelang/runtime wrap arrays (they count reads, for tests).

Views read sources through a block cache (64 KiB blocks, 256 blocks; reads over 8 blocks bypass it). When a view over a source with write commits, the source is updated in place: the patches before the first size change as they are, then everything from that change on in one write, and truncate(n) (required) if the data shrank.

Views

A view is a Proxy over the bytes. Reading a field scans its struct (scalars are read, each field’s byte range recorded); inline structs and arrays become child views, and pointer and at targets, byte fields, codec regions and stream fields are read on first access.

MemberMeaning
$offset, $sizeThe struct’s position and size in its byte space.
$raw()The struct’s current bytes (in-place edits applied).
$set({ … })Sets several fields.
$patches(options?)The edits as { at, remove, insert } patches against the original bytes.
$commit(options?)Applies the edits, returns the new bytes and rebinds the view to them.
$plain()A deep plain snapshot.
$source(field)A SyncSource over a byte field, stream field or codec region, read lazily.

Inner.view(outer.$source("file")) nests a view over decoded or joined bytes through the same caches.

Edits. Setting a scalar to a value of the same encoded size that no derived or computed field depends on patches the bytes in place. Anything else (array mutators, size changes, fields that later reads depend on, replacing a struct or target) materializes the node; $commit then re-encodes the root with the generated writer, passing unchanged byte arrays through, reusing the encoded bytes of unchanged codec regions, and keeping placeables where they were. The result is diffed into patches. $commit({ rewrite: true }) takes this path without edits.

Grow policies ($commit({ grow })):

  • default: targets that grow move to the end; everything else stays unless growth before it pushes it forward (by the least that keeps its alignment);
  • "append": also targets pushed by inline growth move to the end;
  • "splice": nothing moves to the end; every size change shifts what follows;
  • "error": any size change is a DERIVE error.

Lazy streams and codecs. A stream (T x from arr[*].data) is a byte store over its pieces: an offset maps to a piece through the pieces’ decoded lengths, taken from the element’s own fields when the schema states them (a counted u8[n], or u8[] inside @size(n)), and otherwise learned by decoding the pieces in order. Only the touched pieces are decoded, and at most 16 pieces (64 MiB) are cached. A @ranged codec region decodes only the 4 KiB blocks a read touches (256 cached). Regions over a stream keep their end unknown until something needs it. A @base struct followed by inline data is measured with the eager reader.

Async views. Every property chain (av.entries[3].data) is awaitable. Awaiting walks the sync view; when it needs bytes that aren’t fetched or an async codec’s output, the walk is abandoned, the data is fetched and the walk retried (scans and decoded regions stay cached). Async views change data with await av.path.$set({ … }), and $commit resolves to the patches (written back to the source when it has write). $read(f) runs a side-effect-free function against the sync view at that path, retrying as needed.

Async parsing. Parsing and serializing are sync unless the format has an async codec; then only parseAsync/serializeAsync and async views exist, and an async view commits structural edits through the async writer.

Host interfaces

Externs are configured per target in yaffle.json (targets.ts.externs): extern name → module (relative paths are relative to yaffle.json), optionally module#export; the default export name is the extern’s name. An unconfigured extern fails with INPUT when used. The extern types uleb128 and sleb128 are built in.

// extern fn pathHash(char path[]) -> u32
export function pathHash(path: string): number { … }

// extern codec zstd / extern async codec oodle / extern codec chacha20(...) @ranged
export const zstd = {
  decode(input: Uint8Array, ctx: { args: unknown[]; outputSize?: number; position?: number }) { … },
  encode(input: Uint8Array, ctx) { … }
};

// extern type VarInt(u8 maxBytes) : u64
export const VarInt = {
  size(bytes: Uint8Array, maxBytes: number): number | undefined { … }, // undefined: need more
  read(bytes: Uint8Array, maxBytes: number): bigint { … },
  write(value: bigint, maxBytes: number): Uint8Array { … }
};
  • Arguments arrive converted to their declared types (64-bit integers as bigint, byte arrays as Uint8Array). Exceptions thrown by host code become CODEC errors.
  • Codecs get ctx.args, ctx.outputSize (when an inner @size gives the decoded size) and, for @ranged codecs, ctx.position (the slice’s offset in the region). Async codecs return Promises. @streaming codecs return { output, consumed }. Ranged codecs must map byte i to byte i and encode deterministically: views decode their blocks independently and re-encode them on commit.
  • Codec reuse: a value read from a codec region remembers the encoded and decoded bytes. When the same object is serialized and its plain encoding is unchanged, the original encoded bytes are written instead of re-encoding, so rebuilds are byte-exact even for codecs that don’t re-encode identically (zstd). @hidden stream pieces stay attached to their struct, so serialize(parse(bytes)) reuses every piece.
  • Extern types with @size(n) skip size; write must return exactly n bytes.

Runtime structure

@yafflelang/runtime has no dependencies; generated modules import it whole.

ModuleContents
errors.tsYaffleError, path helpers, safe/safeAsync.
prim.ts, int.tsPrimitive encodings; integer semantics of IR.md §6 (truncating /, 64-bit bitwise, ranges).
strings.tsText encodings, terminators, code-unit lengths.
checksum.tscrc32, adler32.
arrays.tsmap, filter, indicesWhere, sortedBy, unique, concat.
json.tsCanonical JSON conversions, raw values of one-way transforms, hidden stream pieces.
source.tsSources, byte stores, block caches, NeedBytes/NeedAsync control signals.
reader.tsThe read context: windows, bases, targets with identity per address and type, unions.
writer.ts, layout.tsThe relocation-style writer (blocks, placeables, fixups, regions) and the layout engine of IR.md §4, with sticky layout for view commits.
codec.ts, helpers.tsCodec calls and reuse, stream join/split, write-side checks, LEB128.
api.ts, open.tsGlue for the generated roots: parse/serialize, opening views, view types.
view.ts, lazy.tsLazy struct and array views, byte spaces and patches; piece and ranged-block stores.
commit.ts, async.ts$patches/$commit; async views.

The language service (packages/compiler/src/service) reads views through nodeOf, StructNode/ArrayNode, Lazy.target and Space to annotate sample files.

Tests

npx vitest run --project @yafflelang/runtime
npx vitest run --project yaffle packages/compiler/test/backends/ts
node --max-old-space-size=12000 --conditions=development conformance/src/cli.ts --targets ts
YAFFLE_TS_VIEW_REWRITE=1 node --max-old-space-size=12000 --conditions=development conformance/src/cli.ts --targets ts
  • The backend tests compile IR fixtures (ir.ts) or .yfl sources, typecheck the output with tsc --strict and run it.
  • --targets ts runs both ts (parse, serialize, round trips) and ts:view (every parse case through a view, committed without edits). With YAFFLE_TS_VIEW_REWRITE=1, view commits use { rewrite: true }, exercising the structural-edit path.

Known gaps

  • @split(fit) needs element placements to know each block’s capacity. After a view edit that changes a stream’s length, the pieces’ size fields from the view no longer cover it (DERIVE).
  • Union values inside views are snapshots: editing one materializes the parent. Derived and computed fields of an edited view node are stale until $commit.
  • Async ranged codecs decode their whole region in views; an async view’s structural commit fetches the whole source first.
  • NaN payloads are not preserved (f16 NaN is written as 0x7e00; f32/f64 depend on the engine).
  • A target whose content aligns relative to its base must itself start aligned: @align inside a target that is neither fixed nor aligned at its start is computed as if it were.
  • The @origin of elements of an array of pointers may only use earlier fields and this (elements are followed right away, field pointers after the struct).
  • A checksum in a nested @base that covers a @ref resolved by an enclosing base is computed over placeholder zeros.
  • inlineRuntime (copy the runtime into the output) is a programmatic backend option only.

Rust target

The Rust backend (packages/compiler/src/backends/rust) turns the IR into Rust source that depends on the runtime crate yaffle-runtime (packages/runtime-rust, no dependencies), referred to as yr, or carries a copy of it inline (inlineRuntime).

Generated code

One module file (mod.rs) with every type, its reader, writer and canonical JSON, the root API, and one re-export submodule per IR module. Target options (targets.rust, all optional):

  • crate: { name?, runtimePath? }: emit a whole crate (Cargo.toml + src/lib.rs);
  • runtime: the runtime’s Rust path in module mode (default ::yaffle_runtime);
  • fileName: the module file name (default mod.rs);
  • views: also generate views (X::view, XView);
  • dependencies: crates the externs need ({ "zstd": "0.13" }), added to Cargo.toml;
  • externs: see Host interfaces.

Generated code compiles without warnings, clippy included.

Root API

#![allow(unused)]
fn main() {
#[derive(Debug, Clone, PartialEq, Default)]
pub struct Table { pub count: u32, pub entries: Vec<Entry>, /* … */ }

impl Table {
    pub fn parse(bytes: &[u8]) -> yr::Result<Table>;
    pub fn parse_with(bytes: &[u8], opts: &yr::ParseOptions, params: &TableParams) -> yr::Result<Table>;
    pub fn serialize(&self) -> yr::Result<Vec<u8>>;
    pub fn serialize_with(&self, params: &TableParams) -> yr::Result<Vec<u8>>;
    pub fn to_json(&self) -> yr::Value;
    pub fn from_json(v: &yr::Value) -> yr::Result<Table>;
}
pub struct TableParams { pub endian: yr::Endian, pub version: u32, /* root parameters */ }
}

yr::ParseOptions { strict } turns on strict mode. Errors are yr::YaffleError with code, path and offset (docs/CONFORMANCE.md §4); write errors carry the writer’s path. yr::Value is the runtime’s own JSON type (ordered objects, parser and printer).

Types

One Rust struct per IR struct serves as both parse output and serialize input. Derived, computed, constant and optional-input fields stay in the struct; the writer ignores what the IR says is not an input.

IRRust
u8…u64, i8…i64u8…u64, i8…i64 (u24 → u32, u40–u56 → u64, range-checked)
f16, f32, f64f32, f32, f64
bool, strings, u8[], other arraysbool, String, Vec<u8>, Vec<T>
pointersRc<T>; Option<Rc<T>> when nullable; Rc<Vec<T>> with @nullable(empty)
recursive inline fieldsBox<T>
enumsRust enums (@open adds Unknown(raw))
unionsan enum with one tuple variant per IR variant, named after its first use
switchesenums, with an accessor per field (fn body(&self) -> Option<&T>)
conditional fields, optional inputsOption<T>
one-way transformsyr::OneWay<V, R> (JSON holds the raw value)
extern typestheir value type

Hidden members, which compare equal to anything and default to empty, so hand-built values use Table { entries, ..Default::default() }:

  • __id: yr::Ident on structs that pointers target: one object per (region, address, type), so shared targets and cycles print as $ref JSON Pointers and are written once;
  • __cache: yr::ViaCache on structs with codecs: the original encoded bytes per region, reused when the decoded bytes and arguments are unchanged, so rebuilds are byte-exact;
  • @hidden stream carriers, stored so the stream can be joined and never in JSON.

Integer expressions are computed exactly in i128 (bitwise ops with infinite-precision two’s complement); overflow and out-of-range stores are RANGE.

Runtime structure

packages/runtime-rust/src; modules refer to each other only through super::, so the crate can be inlined as a module.

ModuleContents
read.rsReader: position, base, region end, strict mode, identities, decoded regions, padding
write.rsWriter: relocation-style encoding into units, then the layout engine and fixups
view.rsViewRoot, ViewState, scans, patches, Grow commits
space.rsbyte spaces: plain bytes or a LazyBuf filled block by block (Source, FileSource)
codec.rsCodec, ExternType, Arg, CodecCtx, ViaCache, built-in uleb128/sleb128
json.rsValue, canonical JSON (docs/CONFORMANCE.md §3), $ref resolution
int.rsexact integer arithmetic
error.rsYaffleError, ErrorCode, paths
prim.rs, strings.rs, checksum.rs, regex.rs, helpers.rsprimitives, text encodings, crc32/adler32, the in /re/ subset, small helpers for generated code

The writer encodes first and lays out second: each base and placeable becomes a unit of inline bytes, because layout paths (records[*].model) must see placeables found inside other placeables. Position-dependent padding is recorded as pads and sized when the unit is placed; values that depend on layout are fixups run afterwards (values, then @ref, then byte-reading fixups such as checksums in dependency order, then late checks). Nested @base regions and codec regions are finished into bytes when they end. @ref takes the placed candidate with the same identity, else the first equal value; a pointer with no candidate in its base is handed to the enclosing base.

Host interfaces

targets.rust.externs[name] is a Rust path used verbatim (crate::hash::gt_hash) or a file (./externs/zstd.rs#ZSTD, relative to the output directory; the item defaults to the extern’s name). Files are mounted with #[path] as ext_<stem> modules; the conformance adapter copies them into its harness crate.

ExternImplementation
extern codec ca value implementing yr::Codec (decode, encode); arguments arrive as yr::Args in CodecCtx { args, output_size, position }
@streaming codecCodec::decode_streaming, returning the output and the bytes consumed
extern type T(args) : Va value implementing yr::ExternType<Value = V>
extern fn f(…) -> Ra function taking typed arguments (integers as their Rust type, strings &str, bytes &[u8])

conformance/externs/rust/stub/yaffle_runtime.rs mirrors Arg, CodecCtx, Codec and ExternType, so the extern files can be unit-tested on their own (conformance/scripts/test-rust-externs.sh); keep it in sync with codec.rs. Async codecs run synchronously.

Views

With views: true:

#![allow(unused)]
fn main() {
let v = Table::view(bytes)?;               // or view_with(bytes, opts, params), view_source(Box<dyn yr::Source>, …)
let n = v.count()?;                        // getters read on demand
let e = &v.entries_views()?[3];            // views of inline and `at` structs, arrays of them, pointer targets
e.set_handler(7)?;                         // patched in place when possible, else a structural edit
let whole: Table = v.materialize()?;       // the whole value, read through the views
let bytes = v.commit_with(yr::Grow::Append)?; // Splice | Append | Error; commit() returns the patched bytes
}

Also loc(), shallow(), span(field), refresh(), patches(), params(), and replace_<field> for at fields and owned pointers (the new target is appended).

  • One yr::ViewRoot per document holds the bytes, the reader state every read shares (identities, decoded regions, streams) and a node cache: every handle of the struct at (region, position, type) shares one ViewState.
  • Scans read inline members only; pointer and at targets, from fields and non-streaming codec regions are read when asked for. @ranged codec regions decode block by block, file sources load in 64 KiB blocks, and from arr[*].field streams fill piece by piece.
  • set_<field> patches in place when the field was scanned, keeps its encoded size, and neither a derived or computed field nor a layout item argument reads it; otherwise it records a structural edit.
  • commit_with: Splice re-serializes the materialized value. Append keeps every placeable where it was when it still fits and is still aligned, and puts the rest at the end; it falls back to Splice when the original document doesn’t rebuild byte-exact. Error is Append that fails if anything moves or changes size.

Tests

export CARGO_TARGET_DIR=/tmp/yaffle-cargo-target
# Runtime unit tests and lints
cd packages/runtime-rust && nix shell nixpkgs#cargo nixpkgs#rustc nixpkgs#clippy \
  --command sh -c 'cargo test && cargo clippy --all-targets'
# Backend tests: generate, build and run crates per test file
npx vitest run --project yaffle packages/compiler/test/backends/rust
# Conformance, parse and view mode (rust and rust:view)
node --max-old-space-size=12000 --conditions=development conformance/src/cli.ts --targets rust
# Extern unit tests against the stub runtime
nix shell nixpkgs#rustc --command conformance/scripts/test-rust-externs.sh

toolchain.ts uses cargo from PATH, else nix shell nixpkgs#cargo nixpkgs#rustc. Builds share CARGO_TARGET_DIR (default $TMPDIR/yaffle-cargo-target); YAFFLE_RUST_RUNTIME_DIR points them at another copy of the runtime crate, and test crates live in $TMPDIR/yaffle-rust-tests/<name> (YAFFLE_RUST_TEST_DIR). Each vitest file builds one crate with its fixtures and a harness binary that runs JSON operations from stdin (harness.ts).

Known gaps

  • No async API: a Source is read on demand (blocking) and async codecs run synchronously.
  • @align(from: file) inside a @base whose position isn’t known while it is written (a @base inside a pointer target) is a DERIVE error.
  • @origin that measures fields is supported on plain pointer fields only (the target is followed after the struct’s members); elsewhere it is a generation error.
  • @ref escalation stops at codec regions (REF); checksums of a nested base that cover an escalated pointer see its bytes before they are filled.
  • @split(fit) assumes an element’s encoded size grows with its piece (binary search).
  • Layout-dependent values inside codec regions that refer to fields outside them, and layout-dependent bitfields, are rejected at generation time.
  • Views: unions and fields read through @origin pointers in arrays have no child views; edits inside decoded regions and streams are structural; handles other than the root are stale after a structural commit; Append keys kept positions by writer path, so inserting array elements moves more than needed.
  • replace_<field> applies only the placement alignment and the struct’s own layout item for that field (not an enclosing struct’s layout, nor from/offset), and isn’t generated when an alignment needs the writer ($index, offsets, measurements).

Python target

The Python backend (packages/compiler/src/backends/python) turns the IR into a Python package with one module per IR module, over the pure-Python runtime yaffle_runtime (packages/runtime-python, Python 3.12+, no dependencies). Generated code and the runtime are mypy --strict clean.

Generated API

from gen.archive import Archive, Entry

pf = Archive.parse(data)                          # bytes, bytearray, memoryview or a sync source
pf.entries[3].align = 0x200                          # plain dataclasses
out = pf.serialize()                                 # derived fields (count, offset, size) recomputed
r = Archive.safe_parse(data)                      # SafeResult(ok, value, error)
j = pf.to_json(); pf2 = Archive.from_json(j)      # canonical JSON, $ref for shared targets
v = await Node.parse_async(src, key=k, nonce=n)      # root parameters are keyword-only arguments
new = Archive(entries=[Entry(handler=0, kind=1, align=0x100, data=b"..")])   # inputs only

v = Archive.view(rt.file_source("big.bin"))     # lazy, memory-mapped
v.entries[3].handler = 7                             # same size: patched in place
v.entries.append(Entry(...))                         # structural: re-encoded on commit
v._patches(); v._commit(grow="append")               # "default" | "append" | "splice" | "error"
  • Root API, on the root’s dataclass: parse, safe_parse, parse_async, safe_parse_async, from_json, options_from_json, view, view_async (class methods) and serialize, serialize_async, to_json (instance methods). Roots that reach an async codec only get the async parse and serialize methods. Keyword arguments: endian ("le"), strict (parse and views) and every root parameter (bound parameters without $); parameters with defaults are optional.
  • Modules: IR module formats/archive becomes formats/archive.py, with an __init__.py per package. Modules import each other as module objects (relative imports), so cycles between modules work. With inlineRuntime the runtime is copied into _yaffle_runtime/ and imported relatively; targets.python.runtimeModule names another runtime module.

Types

yafflePython
integers of every widthint (exact)
f16 / f32 / f64float (the exact double of the stored value)
bool, strings, u8[]bool, str, bytes
other arrayslist[T]
struct@dataclass(kw_only=True)
enumenum.IntEnum (open enums: E | int)
unionyaffle_runtime.Variant(type, value)
nullable pointerT | None
one-way transformthe transformed value (canonical JSON shows the raw value)
  • Dataclass fields are in declaration order (blocks, switches and bits flattened). Inputs the caller must give have no default; derived and computed fields, constants, optional inputs and fields of conditional branches have defaults (0, "", the constant, None, …), so X(...) takes only the inputs. A field declared in several branches gets the union of its types.
  • Absent fields: readers and from_json create objects with cls.__new__ and only assign the fields they read, so an untaken branch’s field is absent (the class default shows through) and a missing required input is reported as INPUT with the field’s path.
  • Names: structs keep their IR names; a trailing _ avoids keywords and builtins. Field attributes are the yaffle names with a trailing _ for hard keywords, builtin type names and, on root classes, the API method names; they never start with _. JSON keys are always the yaffle names. Enum members avoid keywords (None_). An exported root whose name differs from its class gets a module-level alias.
  • Values in: writers accept an enum member, its name or (open enums, known values) an int; True/False are rejected for integer fields and ints for bool fields.
  • Side data (raw values of one-way transforms, @hidden stream pieces) lives in the instance __dict__ under _yfl_raw / _yfl_hidden, so == and repr ignore it.

Runtime structure

ModuleContents
errorsYaffleError (code, path, offset, causes), path prefixes while unwinding
primprimitive encodings and range checks
intstruncating division, IEEE float division, shifts, rounding, lengths (IR.md §6)
stringstext encodings, code units
checksumcrc32, adler32
arraysaggregates, derived arrays (map, filter, sortedBy, …), value equality
jsonccanonical JSON ($refs, hex, big integers), Variant, side data
valuesenums, bitfields, input validation, after-layout checks
externsloading and calling host code, built-in LEB128 types
codec@via codec calls and codec reuse
streamsstreams outside views: joining and slicing, splitting on write
readerReader: the read context, targets with identity, unions
writerWriter, regions, blocks, placeables, fixups, @ref resolution
layoutthe layout engine, including sticky layout for view commits
apiwhat the root API runs on; sources, file_source
storebyte stores filled on demand (block cache, joined streams, ranged regions)
viewstruct and array views, the scan helpers generated scanners call
commit_commit / _patches: fast in-place path, re-encoding with provenance, diff
async_viewawaitable chains over async sources and async codecs
_harnessthe conformance harness (not copied by inlineRuntime)

Readers are imperative functions per struct over Reader; the writer is relocation style (blocks, placeables, fixups, extents for sizeof/offsetof/bytesof/encodedSize). The layout engine, @ref resolution and the commit planner follow the TS runtime, so both targets make the same placement decisions.

Host interfaces

targets.python.externs maps an extern to a .py file (relative to the yaffle.json, or absolute) or an importable module name, optionally with #Item (default: the extern’s name). Files are loaded by path, so extern files need not be packages.

ExternPython object
extern fn f(T a, …) -> Ra callable f(a, …); ints are int, u8[] bytes, char[] str
extern codec c(args)decode(data, ctx) -> bytes and encode(data, ctx) -> bytes; ctx is CodecContext(args, output_size, position)
@streaming codecdecode returns (output, consumed)
@ranged codecmay be called on any slice, with ctx.position
async codecdecode / encode may return awaitables
extern type X(args) : Tsize(data, *args) -> int | None, read(data, *args), write(value, *args) -> bytes

Any exception from host code is a CODEC error. An extern without an implementation fails with INPUT when used.

Views and async

  • Views are lazy: a struct view (StructView) scans its scalars and records field ranges; inline structs and arrays are child views, and targets, bytes, codec regions and streams are read on access. View methods start with _ (_commit, _patches, _offset, _size, _raw, _set, _fields, _plain), like namedtuple’s. Array views are mutable sequences. A struct view works as the input of X.serialize(view) and X.to_json(view).
  • Sources: bytes are viewed in memory; file_source(path) is memory-mapped (in-place patches go to a copy-on-write map); other sync sources are read through a block cache (64 KiB × 256). @ranged codec regions decode only the 4 KiB blocks a read touches, and from arr[*].field streams decode only the pieces it touches.
  • Edits: a same-size scalar edit that no derived or computed field depends on is patched in place; anything else marks the view structural, and _commit re-encodes with the generated writer: unchanged bytes keep their provenance, unchanged codec regions reuse their encoded bytes, and placeables keep their positions where the grow policy allows. Writable sources (file_source(path, writable=True), or any source with write/truncate) get the commit written back.
  • Async: parse_async, serialize_async and safe_parse_async accept async sources (a read returning an awaitable) and async codecs. X.view(async_source), X.view_async(any) and the views of async roots are async views: attribute and item chains are awaitable and resolve to plain snapshots (dicts, lists); edits go through await chain._set(...) and await view._commit(). Missing bytes and async codec output are fetched and the walk retried.

Tests

# conformance (python and python:view); Python comes from $YAFFLE_PYTHON, python3 on PATH or
# `nix shell nixpkgs#python3`
node --max-old-space-size=12000 --conditions=development \
  conformance/src/cli.ts --targets python
# backend tests (vitest), including the Python unit tests; YAFFLE_PYTHON_MYPY=1 adds mypy --strict
YAFFLE_PYTHON_MYPY=1 npx vitest run --project yaffle packages/compiler/test/backends/python
# Python unit tests directly
PYTHONPATH=packages/runtime-python python3 -m unittest discover -s packages/runtime-python/tests
python3 -m unittest discover -s conformance/externs/python

Caches live outside the repository, in $YAFFLE_PYTHON_CACHE (default <tmpdir>/yaffle-python-cache): the resolved interpreter, bytecode, generated test packages and mypy’s cache. The conformance adapter runs one Python process per case directory, which writes one result file per case.

Known gaps

  • Union values in views are snapshots (editing one replaces the value). A sync view of a union that reaches an async codec raises INPUT (“not supported”).
  • Bytes fields in views are read whole on first access.
  • Async views resolve to plain snapshots, not dataclasses.
  • A stream whose piece lengths the schema doesn’t state decodes all its pieces once to learn its size.
  • f16/f32 NaN payloads go through C float conversions and may be quieted; "NaN" in JSON writes the canonical quiet NaN.
  • Python bool is an int: writers tell them apart, but plain == in user code does not.
  • Elements of an array of pointers with @origin are followed right away, so their origin may only use earlier fields.

Go target

The Go backend (packages/compiler/src/backends/go) turns a yaffle program into one Go package; the generated code imports the runtime yafflelang.org/runtime as yaffle (packages/runtime-go, Go 1.22, no dependencies). Every IR module becomes one .go file of the package: yaffle allows type cycles across modules, Go packages can’t import each other in cycles.

The runtime module lives in the packages/runtime-go subdirectory of the repository. The page yafflelang.org/runtime (docs/runtime/index.html) declares it with a go-import tag carrying that subdirectory, a form go get understands from Go 1.25; release tags are prefixed with it (packages/runtime-go/v0.0.1).

Generated API

For every exported struct X:

type X struct { … }                     // one type for output and input
type XOptions struct {                  // root parameters, endian, strict
	Endian  yaffle.Endian
	Strict  bool
	Version uint32                      // a root parameter; with a default: *uint32 (nil = default)
}
func ParseX(data []byte, opts *XOptions) (*X, error)
func ParseXContext(ctx context.Context, data []byte, opts *XOptions) (*X, error)
func SerializeX(v *X, opts *XOptions) ([]byte, error)
func SerializeXContext(ctx context.Context, v *X, opts *XOptions) ([]byte, error)
func ViewX(data []byte, opts *XOptions) (*XView, error)    // see "Views"
func XToJSON(v *X) ([]byte, error)      // canonical JSON (docs/CONFORMANCE.md §3), $refs
func XFromJSON(data []byte) (*X, error)
func XOptionsFromJSON(data []byte) (*XOptions, error)

opts may be nil. Errors are *yaffle.Error (Code, Path []any, Offset, Detail, Causes); yaffle.AsError(err) unwraps them.

Target options (targets.go in yaffle.json, or targetOptions): package (default gen), runtimeModule (the runtime’s import path; required with inlineRuntime, which copies the runtime’s sources into yaffle/), views (default true), materialize (adds MaterializeXView, used by the conformance view mode), externs and dependencies (below).

Type mapping

yaffleGo
u8/u16/u32/u64, i…uint8…uint64, int8…int64
u24, i24 / u40…u56, i40…i56uint32, int32 / uint64, int64 (range-checked on write)
f16 / f32 / f64float32 (exact) / float32 / float64
bool, char/str…bool, string (UTF-8; other encodings are converted)
u8[n] (bytes)[]byte (aliases the input after a parse; decoded regions are copies)
arrays[]T
structs*S (pointers: shared targets and cycles are the same object)
enumstype E uint16 with constants EMember, IsKnown(), String(); open enums hold any value
unions*AOrB with Type string and one field per variant
pointersthe target’s type; nullable scalars and strings *T, nullable structs and slices nil
conditional fields, optional inputs*T for scalars and strings, nil slices and pointers
one-way transforms (as without an inverse)the transformed value F plus the raw value FRaw (written, and shown in JSON)
@hidden fieldsunexported fields
extern typestheir value type

Derived, computed and constant fields stay in the struct (filled by a parse, recomputed by a write). A nil slice means “absent” only for conditional fields; elsewhere it is an empty slice. Every struct carries an unexported yfl yaffle.Meta (codec reuse, lazy stream pieces).

Host interfaces

Externs are configured in targets.go.externs. A spec is a file path (optionally #Item) or importpath.Name:

  • File: the file becomes part of the generated package (its package clause is rewritten), so the item may be unexported. Each extern file is copied on its own, so it must be self-contained, and helper names must not collide between the extern files of one package.
  • importpath.Name: imported with an alias (ext1, …). Extra modules go into targets.go.dependencies ({ "github.com/klauspost/compress": "v1.20.1" }).
  • Missing: a placeholder is generated; using it fails with INPUT.
// extern codec c(args)
type Codec interface {
	Decode(input []byte, ctx *yaffle.CodecCtx) ([]byte, error)
	Encode(input []byte, ctx *yaffle.CodecCtx) ([]byte, error)
}
// @streaming: Decode also reports the bytes consumed
type StreamingCodec interface {
	Decode(input []byte, ctx *yaffle.CodecCtx) (out []byte, consumed int, err error)
	Encode(input []byte, ctx *yaffle.CodecCtx) ([]byte, error)
}
// CodecCtx: Context (from ParseXContext/SerializeXContext), Args []any (u8 → uint8, u64 → uint64,
// u8[n] → []byte, …), OutputSize (-1: unknown), Position (@ranged).

// extern type T(args) : V
type ExternType[V any] interface {
	Size(data []byte, args []any) (n int, ok bool, err error)  // ok=false: need more bytes
	Read(data []byte, args []any) (V, error)
	Write(value V, args []any) ([]byte, error)
}

// extern fn f(u8 data[]) -> u32
func f(data []byte) uint32   // plain Go types

yaffle.ArgInt, ArgUint and ArgBytes read arguments. Every host failure (an error or a panic) is a CODEC error. extern async codecs are ordinary blocking codecs: a blocking call on a goroutine is Go’s equivalent of async, and the context of the …Context entry points reaches codecs through CodecCtx.Context, so a slow codec can honour cancellation. The built-in uleb128/sleb128 are yaffle.Uleb128 and yaffle.Sleb128.

Views

Generated for every struct unless targets.go.views is false: lazy reads, in-place patches, and structural edits committed through the writer with a sticky layout and a diff.

v, err := gen.ViewArchive(data, nil)    // *ArchiveView (embeds *yaffle.Node)
n, err := v.Count()                        // scalars come from a shallow scan
es, err := v.EntriesViews()                // views of inline structs: no further reads
d, err := es[1].Data()                     // a target is read alone, on demand
err = es[1].SetHandler(2)                  // same size, nothing depends on it: patched in place
err = es[1].SetData(bigger)                // structural: applied at commit
all, err := v.Entries()                    // whole values when a getter needs more
p, err := v.Patches(yaffle.CommitOptions{Grow: "append"}) // []yaffle.Patch{At, Remove, Insert}
out, err := v.Commit(yaffle.CommitOptions{})              // new bytes; the root view rebinds
  • Scans. A view runs the generated reader in a shallow mode: pointer and at targets, from fields nothing else reads and lazy codec regions are not read; the scan records where they are (with a closure that replays the read in the reader state it was found in), every field’s byte span and each struct’s arguments, per struct value. Inline children get views from the same scan. A nested @base struct, and a struct whose read-side expressions read targets, are read with targets followed.
  • Getters return values from the scan when they are complete there, read a target or a deferred field alone, or else read the node completely. FView()/FViews() return views of inline structs, targets, deferred fields and arrays of them. Target values are cached per (address, identity key), so shared targets stay shared.
  • Setters encode the value with the generated writer code for that field (checks run immediately). If the encoding has the field’s size and nothing reads, derives from or lays out by the field, the bytes are patched in place; anything else is a structural edit.
  • Commit. In-place patches alone (and no ancestor derives from child content) are returned as they are. Otherwise the document is read with origins tracked, the edits are applied, the root is re-serialized with the sticky layout (grow policies default, append, splice, error), codec regions reuse their bytes, and the output is diffed against the original, anchored on byte slices that still alias it.
  • Lazy stores. A stream joined from pieces (from arr[*].data) is a PieceStore: pieces are decoded on demand (their lengths come from the data where the schema states them), with an LRU of 16 pieces and 64 MiB. A @ranged codec region is a RangedStore of 4 KiB blocks decoded alone (256 blocks cached). Readers over a store keep a window of it.

Runtime structure

Generated readers and writers panic with *yaffle.Error; the entry points recover it and return it, so host code only sees errors. A shared PathStack records struct frames and array indices as code runs and isn’t unwound by the panic, so the entry point attaches the exact path.

FileContents
api.gothe entry points’ glue (ParseWith, SerializeWith, ToJSONWith, FromJSONWith)
errors.goError, codes, the path stack, Catch/Try
prim.gobyte orders, primitive kinds, raw integers, half floats, range checks
expr.goexpression semantics: exact integers (int64, uint64, *big.Int), lengths, builtins
derived.goderived arrays (map, filter, sortedBy, …) and value equality
strings.gotext encodings
reader.goReader: primitives, padding, targets with identity per (address, type key), unions
writer.goWriter: primitives, padding, fixups, targets, nested regions, write-side checks
region.goblocks, placeables, regions: assembly, fixup phases, @ref resolution
layout.gothe layout engine (items, groups, pull-out, alignments, fixed placements, sticky mode)
codec.gocodec calls and codec reuse (Meta)
extern.gohost interfaces, extern types and functions, placeholders, LEB128
stream.gostreams on read and write, byte stores
json.gocanonical JSON writer and reader
scan.goview scans and the reader’s view hooks
views.godocuments and nodes, getters, setters, commit, diff
commit.gothe writer’s commit hooks (origins for the sticky layout)
materialize.goreading a document through its views (conformance view mode)

The generator: index.ts (files, options), model.ts (Go names and types, static sizes, placeables, type keys), irutil.ts (pure IR analyses), expr.ts (expressions with interval analysis: int64 or uint64 when the ranges provably fit, *big.Int helpers otherwise), emitter.ts (what the reader and writer emitters share), read.ts, write.ts (relocation-style: derived values that need sizes or positions are fixups), json.ts, decls.ts, api.ts, view.ts with viewinfo.ts (what a scan reads, skips and defers), module.ts (imports, hoisted variables, externs), code.ts (code builder, identifiers, unused temporaries) and gofmt.ts (go/printer’s operator spacing and line breaking, applied to the generated subset of Go, so the output is gofmt-stable without running gofmt). conformance.ts and harness.ts are the conformance adapter.

Tests

# conformance in parse and view mode (Go from PATH, $YAFFLE_GO, or `nix shell nixpkgs#go`)
node --max-old-space-size=12000 --conditions=development conformance/src/cli.ts --targets go
# backend tests (skip without a Go toolchain); format.test.ts checks that generated code is gofmt-stable and vet-clean
npx vitest run --project yaffle packages/compiler/test/backends/go
# runtime and extern packages
(cd packages/runtime-go && go test ./... && go vet ./...)
(cd conformance/externs/go && go test ./...)

The adapter runs go with GOTOOLCHAIN=local GOFLAGS=-mod=mod GOSUMDB=off and keeps GOPATH/GOCACHE/GOMODCACHE outside the repository: from the environment when set, otherwise under $YAFFLE_GO_CACHE (default <tmpdir>/yaffle-go-cache). The backend tests write their modules under $YAFFLE_GO_TEST_DIR (default the system temp directory).

Known gaps

  • Rejected at generation (“not implemented”): a field declared in several branches with different Go types, union-valued transforms, and layout item arguments that read sizes, positions or encodings.
  • from: file alignment of a block whose position is only known after layout aligns the block itself; it is exact when the enclosing positions are known or aligned.
  • Views: arrays are read whole by a scan, so a stream or ranged region whose struct has a large inline array is read completely when its view opens; pieces without a stated length are decoded up front to learn the stream’s size. Values inside streams and codec regions can’t be edited (INPUT: set the enclosing field), union variants have no view getters, structs read through a codec of their own have no views unless deferred, and after a commit only the root view is rebound.

C# target

The C# backend (packages/compiler/src/backends/csharp) turns the IR into C# 13 source for .NET 9; the generated code depends on the runtime library Yaffle.Runtime (packages/runtime-csharp, no dependencies), or carries it inline (inlineRuntime).

Generated code

One .cs file per IR module (namespace targets.csharp.namespace, default Yaffle.Generated) and YaffleExterns.cs. Nullable reference types are on; the output builds warning-free at every warning level.

  • One class per struct, public sealed partial class X, used both as the parse result and as the serialize input: Serialize(Parse(bytes)) and Serialize(FromJson(ToJson(x))) work directly. Fields are PascalCase properties (entry_count → EntryCount; collisions get a _2 suffix; the class name and the root API names are reserved). Derived, computed and constant fields are properties too: Parse sets them, Serialize recomputes them.
  • Hidden fields (@hidden) are internal members named Hidden_x.
  • One-way transforms (as without an inverse): the property holds the transformed value and a public XxxRaw property the raw value, which Parse and FromJson set, Serialize writes and canonical JSON shows.
  • Fields of different types in different branches are object?.
  • Readers, writers, JSON converters and view metadata are internal static members of the partial class YaffleImpl (Read_X, ReadAsync_X, Scan_X, Write_X, WriteAsync_X, ToJson_X, FromJson_X, Info_X, enum helpers).

Type mapping

IRC#
u8…u64, i8…i64byte…ulong, sbyte…long (u24 → uint, u40–u56 → ulong)
f16, f32, f64Half, float, double (NaN is written canonically)
bool, stringsbool, string
u8[], other arraysbyte[], List<T>
enumsC# enums over the storage type (open enums keep any value)
unionsRt.UnionValue { Type, Value }
conditional fields, nullable pointers, ?= inputsnullable (uint?, Entry?)
extern typestheir value type

Required reference-typed inputs are non-nullable (= null!); Serialize and FromJson report a missing one as INPUT with the field’s path. Closed enums reject non-member values on write (CHECK). Integer expressions are computed exactly in long, Int128 or BigInteger (chosen by interval analysis) and land in typed locations through range-checked conversions (RANGE).

Root API

Static members of each root’s class, with an options class XOptions : Rt.ParseOptions (Endian, Strict, the root parameters by name, and XOptions.FromJson for canonical-JSON args):

X Parse(byte[] | Rt.ISource | Stream data, XOptions? options = null)
Rt.SafeResult<X> SafeParse(byte[] data, XOptions? options = null)
ValueTask<X> ParseAsync(byte[] | Rt.IAsyncSource data, XOptions? options = null)
ValueTask<Rt.SafeResult<X>> SafeParseAsync(byte[] data, XOptions? options = null)
byte[] Serialize(X value, XOptions? options = null)
ValueTask<byte[]> SerializeAsync(X value, XOptions? options = null)
JsonNode ToJson(X value)
X FromJson(JsonNode? json)
Rt.View<X> View(byte[] | Rt.ISource data, XOptions? options = null)
Rt.AsyncView<X> ViewAsync(byte[] | Rt.IAsyncSource data, XOptions? options = null)

Roots that reach an async codec get only the async members (and ViewAsync). Errors are Rt.YaffleException with Code, Path (field names and long indices) and Offset; generated methods add path segments in exception filters, so nothing is caught and rethrown.

Host interfaces

targets.csharp.externs[name] is path.cs, path.cs#Fully.Qualified.Item, or an item name without a file. Without #Item the item is the extern’s name in PascalCase (Externs.F for functions). Packages an extern needs go in targets.csharp.dependencies.

ExternImplementation
extern codec ca class with a parameterless constructor implementing Rt.ICodec: Decode/Encode(ReadOnlyMemory<byte>, CodecContext) → byte[]; async codecs override DecodeAsync/EncodeAsync
@streaming codecRt.IStreamingCodec: DecodeStreaming(input, ctx) → CodecResult(Output, Consumed)
extern type T(args) : VRt.IExternType<V>: int? Size(bytes, args) (null = need more bytes), V Read(bytes, args), byte[] Write(V, args)
extern fn f(…) -> Ra static method taking the parameters’ C# types

CodecContext carries the arguments (Args, Arg<T>(i)), the decoded size an inner @size gives (OutputSize) and, for @ranged codecs, the block’s position (Position). Host failures become CODEC errors. Sources are Rt.ISource / Rt.IAsyncSource (Size, Read(offset, length)); a source that is also an Rt.ISink / Rt.IAsyncSink receives committed bytes.

Views and async

A view’s root is an object of the regular class, filled lazily:

var view = Archive.View(bytes);              // or View(ISource): block-cached, reads what you touch
view.Root.Entries[3].Align = 0x200;             // plain property and list edits
IReadOnlyList<Rt.Patch> p = view.Patches();     // { At, Remove, Insert } against the original
byte[] updated = view.Commit(new Rt.CommitOptions { Grow = Rt.Grow.Splice });

var av = Archive.ViewAsync(asyncSource);     // async sources and formats with async codecs
var size = await av.GetAsync(r => r.Entries[1].Size);
await av.SetAsync(r => r.Entries[2].Handler = 5);
await av.CommitAsync();
  • Scan_X reads a struct’s scalars and records each field’s byte range and value in an Rt.ViewNode; inline structs become child nodes. Targets, byte arrays, arrays of fixed-size elements, codec regions and from streams are deferred in an Rt.Slot<T> and read on first access. A from stream reads through a store that opens (decodes) a piece only when a read reaches it; @ranged regions decode in 16 KiB blocks.
  • Commit: same-size edits of scalars with an in-place encoder become patches without reading anything else. Any other change re-serializes the root with a sticky layout (placeables keep their positions where the grow policy allows: Default, Append, Splice, Error), unchanged bytes keep their provenance, codec regions reuse their encoded bytes, and the result is diffed into patches. CommitOptions.Rewrite forces the structural path.
  • Async views run every access through a retry loop: a missing source block raises NeedBytes, an async codec NeedAsync; both are awaited and the access re-runs. Formats with async codecs commit structural edits through ViewSpec.WriteAsync.
  • After Commit, the view is rebound to the committed bytes: use view.Root again rather than old references.

Runtime (packages/runtime-csharp/src/Yaffle.Runtime)

FileContents
Api.csParseOptions, source and sink interfaces, array and file sources
Errors.csYaffleException, ErrorCode, Err, SafeResult<T>, control signals
Num.csprimitives, exact integer helpers (Num), range checks (Ck), binary16, Values
Strings.cs, Checksum.cstext encodings and code-unit counts; CRC-32, Adler-32, hex
Arrays.csderived arrays (map, filter, indicesWhere, sortedBy, unique, concat)
Reader.csthe read context: primitives, counts, strings, padding, windows, targets with identity, unions, extern types
Region.cs, Writer.cs, Layout.csthe relocation-style writer: blocks, placeables, fixups and base regions (@ref resolution), the writer generated code calls, and the layout engine of docs/IR.md §4 with sticky layout
Codec.cscodec interfaces and calls, codec reuse, streams, LEB128
Helpers.cswrite-side checks, stream splitting
Json.cscanonical JSON ($ref sharing, a writer without length limits)
Stores.cs, Streams.csbyte stores for views: arrays, block-cached sources, windows, ranged codec regions, joined pieces
ViewState.cs, ViewRoot.cs, Views.cs, View.csview metadata and nodes; scanning, commit planning and patches; View<T> and AsyncView<T>

Tests

.NET comes from $YAFFLE_DOTNET, dotnet on PATH, or nix shell nixpkgs#dotnet-sdk_9 (the resolved path is cached). Caches and build output go to $YAFFLE_CSHARP_CACHE (default $TMPDIR/yaffle-csharp-cache); runtime builds put bin/obj there through $YAFFLE_CSHARP_ARTIFACTS.

# Conformance, parse and view mode (csharp and csharp:view)
node --max-old-space-size=12000 --conditions=development \
  conformance/src/cli.ts --targets csharp
# Backend tests: every schema builds warning-free; the API, views and features end to end
npx vitest run --project yaffle packages/compiler/test/backends/csharp
# Runtime unit tests
nix shell nixpkgs#dotnet-sdk_9 --command dotnet test packages/runtime-csharp/test/Yaffle.Runtime.Tests

The conformance adapter (conformance.ts) builds one long-running host (harness/Host.cs, harness/HarnessCore.cs, the runtime and Roslyn), cached by a hash of its sources. Per case directory it writes the generated code and a small generated harness, and the host compiles them in memory, runs the cases and writes one result file per case.

Known gaps

  • One-way transforms on array elements are read but not written (“not implemented”).
  • Layout item arguments that depend on sizes or positions (sizeof, offsetof, derives through them) are “not implemented”; values, parameters and $index work.
  • Views of unions that reach an async codec are “not implemented”.
  • Views read some values completely during the scan: unions and streaming-codec regions (a view holding one commits through re-serialization), element-placement arrays (T x[] at f($index)), and stream pieces whose decoded length the schema doesn’t give. A deferred array loads all its elements on first access (their own targets stay lazy). @ranged async codecs decode whole regions.
  • In views of a nested (non-root) @base struct followed by more inline fields, the base’s extent covers its inline bytes only.

Java target

The Java backend (packages/compiler/src/backends/java) turns the IR into Java 21 sources that use the runtime in packages/runtime-java (package org.yafflelang.runtime, no dependencies; Maven coordinates org.yafflelang:yaffle-runtime). Generated code and the runtime compile with javac -Xlint:all -Werror.

Generated API

Every struct, enum and union type becomes one file in one package (targets.java.package, default yaffle.gen). An exported struct also gets the root API:

// export struct Doc(u32 version) { … }  →  yaffle/gen/Doc.java
Doc d = Doc.parse(bytes, new Doc.Options().version(2));     // Options: endian, strict, root params
byte[] out = Doc.serialize(d, new Doc.Options().version(2));
SafeResult<Doc> r = Doc.safeParse(bytes, options);          // Ok(value) | Err(YaffleException)
CompletableFuture<Doc> f = Doc.parseAsync(bytes, options);  // also serializeAsync, parseAsync(AsyncSource)
Object json = Doc.toJson(d);                                // Map/List/String/Long/Double/Boolean/null
Doc back = Doc.fromJson(json);
Doc.Options o = Doc.Options.fromJson(argsJson);

Doc.View v = Doc.view(bytes);                               // also view(Source, options)
v.getEntries().get(3).setAlign(0x200);                      // same size: patched in place
v.getEntries().add(Entry.View.of(newEntry));                // structural edit
List<Patch> ps = v.patches();                               // [{at, remove, insert}]
byte[] committed = v.commit(Layout.Grow.SPLICE);            // DEFAULT | APPEND | SPLICE | ERROR

AsyncView<Doc.View> av = Doc.viewAsync(AsyncSource.of(asyncChannel), options);
long size = av.get(w -> w.getEntries().get(3).getSize()).join();
av.set(w -> w.getEntries().get(2).setHandler(5)).join();
byte[] out2 = av.commit().join();

Schemas with an async codec get parseAsync and serializeAsync instead of parse, safeParse and serialize. parse and serialize without options use the defaults. Errors are YaffleExceptions with a stable ErrorCode, a path (field names and indices) and a byte offset (-1 on write).

Struct classes are public final with public fields: one class is both the parse result and the serialize input (derived and computed fields are filled by parse and ignored by serialize). They implement YValue (toJsonValue(), what == on composite values compares) and have a nested View class. @hidden fields are package-private $name fields.

Type mapping

yaffleJava
u8 u16 u24 i8 i16 i24 i32int
u32 u40 u48 u56 i40 i48 i56 i64long
u64long holding the unsigned bit pattern (JSON and expressions treat it as unsigned)
f16, f64 / f32double / float
bool, stringsboolean, String
u8 arrays, other arraysbyte[], List<T> (boxed elements)
pointersthe target type (boxed when nullable)
closed enumJava enum implementing YEnum (value(), member(), fromValue, fromMember)
open enumrecord E(long value, String member) with constants and E.of(v)
union A | Bsealed interface AOrB with one record per variant (named after the variants)
conditional fields, optional inputsboxed and nullable
one-way transformthe value in name, the raw value (written, and shown in JSON) in nameRaw

Switch cases flatten into the struct’s fields (nullable), like the canonical JSON. A field with different Java types in different branches is an Object. A conditional nullable pointer has a presence flag, so a taken branch with a null pointer is null in JSON and an untaken one has no key.

Names: keywords get a _ suffix (open-enum constants named value/member too). User types keep their names; generated code refers to JDK, java.lang and runtime classes by simple name only when no user type shadows them, and fully qualified otherwise. The options class is Options, or Options_ if a user type is called Options.

Generator structure

FileRole
index.tsentry point, extern resolution, inlineRuntime
code.ts, ctx.tsthe indented code builder; per-class imports, hoisted constants, extern references
model.tslookups, Java names and types, field tables, static sizes, placeables, type keys
analysis.tsIR traversals (forEachMember, read-side expressions, extents, $index use, layouts)
expr.tsexpressions with interval analysis: long when the result provably fits, else BigInteger
emitter.tswhat the reader and writer share (contexts, arguments, fills, if/switch)
read.ts, write.tsread/scan and write methods per struct
json.tscanonical JSON (toJson/fromJson)
views.ts, viewclass.tsview analyses, the INFO descriptor and the View class
classes.tsassembles struct, enum and union files and the root API
toolchain.ts, conformance.ts, harness.tsJDK and javac helpers, the conformance adapter and its Java harness

Java lambdas only capture effectively final locals, so generated code keeps mutable state in objects: the error path in a Frame, read-side extents in an int[], write-side extents in Extent objects created up front.

Runtime structure

AreaClasses
readingReader (bounds, windows, padding, strings, identity of targets, unions), Bytes, Encoding
writingWriter (relocation writer), Block, Region, Placeable, Extent, Fixup
layoutLayout (default order, layout { … } items, groups, fixed placements, sticky layout), LayoutItem, LStruct, LArray
valuesInts (integer semantics), Lists (derived arrays), Prim, Rt (helpers for generated code)
codecs, streamsCodecs (host calls, codec reuse), Streams (join, slice, split, slice carriers), Leb128
JSONCanon, Json, Hex, ToJsonCtx, FromJsonCtx
viewsViewRoot, ViewNode, ViewScan, ViewSpace, Store, ListView, StructView, StructInfo, Commit, ViewConv, ViewRead, AsyncView
entry pointsApi, SafeResult, YaffleException, ErrorCode, Source, AsyncSource

Serialization encodes bottom-up into blocks with fixups; each base region is laid out when it is complete, then assembled, then the fixups run. A value decoded through a codec remembers its encoded bytes, so serializing it unchanged reuses them (byte-exact rebuilds of lossy codecs).

Views and async

A generated scan method reads a struct’s scalars into a plain object and records byte ranges; composite fields become view values: child nodes, ListViews, lazy bytes, lazy pointer targets, codec regions decoded on access and lazy streams. Fields that later expressions need are read eagerly. Setters patch the source in place when the new encoding has the same size and nothing depends on the field; anything else is a structural edit. commit() returns the in-place patches directly when it can. Otherwise it turns the view into a plain value (unchanged parts read from the source), re-encodes it with the generated writer under the sticky layout (targets keep their positions where the grow policy allows) and diffs the result into patches.

Byte spaces can be lazy: a Store loads 64 KiB blocks of a Source or AsyncSource, decodes a @ranged codec region block by block, or decodes the pieces of an arr[*].field stream only when a read touches them (when the pieces’ lengths are known from size fields).

Generated readers and writers are synchronous. parseAsync, serializeAsync and the accesses of an AsyncView run on virtual threads; an async codec’s future or a missing block is awaited there, which parks the virtual thread. Accesses of one async view run in order. The typed view (av.view()) can also be used directly; reads then wait on the calling thread.

Host interfaces

ExternHost implementation
extern fn f(T a, …) -> Rpublic static R f(T a, …) (integers as their Java type, u8[] → byte[], char[] → String)
extern codec c(args)a static field implementing Codec: decode(byte[], CodecCtx), encode(byte[], CodecCtx)
@streaming codecalso decodeStreaming(byte[], CodecCtx) returning Decoded(output, consumed)
@ranged codeccalled on any slice of the region; CodecCtx.position() is the slice’s offset
extern async codecAsyncCodec: the same methods returning CompletableFutures
extern type X(args) : TExternType<T>: size(byte[], Object...) (-1: need more bytes), read, write

CodecCtx carries the arguments (boxed: u8 → Integer, u64 → Long bit pattern, u8[n] → byte[]) and outputSize() (-1 when unknown). Host exceptions become CODEC errors.

targets.java.externs in yaffle.json maps extern names to path/file.java[#ITEM] (the file is copied into the output under its public class’s name; ITEM defaults to the extern’s name) or to a class on the classpath, com.example.Host[#ITEM]. Relative paths resolve against targets.java.externsRoot when set. targets.java.dependencies lists Maven jars ({"io.airlift:aircompressor": "2.0.2"}) the conformance adapter and tests download.

The conformance externs are in conformance/externs/java.

Tests

JDK 21 comes from JAVA_HOME, PATH or nix shell nixpkgs#jdk21 (its location is cached). YAFFLE_JAVA_CACHE (default $TMPDIR/yaffle-java-cache) holds the compiled runtime, keyed by a hash of its sources, and downloaded jars; YAFFLE_JAVA_TEST_DIR (default $TMPDIR/yaffle-java-tests) holds the test programs. Use private directories when several checkouts run tests at once.

# conformance, parse and view mode (java, java:view)
node --max-old-space-size=12000 --conditions=development conformance/src/cli.ts --targets java
# backend tests: API, views, lazy and async views, IR features, a view property test,
# -Werror compilation of every conformance schema, and the runtime unit tests
npx vitest run --project yaffle packages/compiler/test/backends/java
# runtime unit tests on their own (a self-contained runner, no JUnit)
nix shell nixpkgs#jdk21 --command packages/runtime-java/build.sh test

The conformance adapter generates the case directory’s program, a harness that runs every case in one JVM, compiles both with -Xlint:all -Werror and judges what the harness observed. In view mode the harness reads every field through the view (ViewRead.materialize) and commits without edits; { file } cases go through viewAsync over an AsynchronousFileChannel.

Known gaps

  • Piece streams whose piece lengths aren’t known from size fields decode every piece on first access.
  • A from stream over a single bytes field forces the whole field (a @ranged source region is still decoded block by block, but all of it).
  • Derived and computed fields of an edited view node are stale until commit(); computed fields can’t read derived fields of array elements.
  • Fields of the same Java type but different JSON rules in different branches (u32 and u64, both long) use the first branch’s JSON rule.
  • In views, edits inside codec regions and streams are always structural, and getters of eagerly read fields return detached views.
  • A layout on the root struct of a codec’s decoded side or a stream is ignored at that level.
  • Unions are named after their variants (AOrB): the IR has no union names.
  • yaffle build passes extern paths relative to the output directory without the directory itself; the backend falls back to the current directory and its ancestors (or targets.java.externsRoot).

C++ target

yaffle build --target cpp generates C++20 for one program: a header with the types and the root API, and a source file with the readers, writers, canonical JSON, views and API definitions. The generated code runs on the runtime in packages/runtime-cpp. The backend lives in packages/compiler/src/backends/cpp.

Building

Options in yaffle.json under targets.cpp:

OptionMeaning
namespaceNamespace of the generated code (default: from the file name)
fileBase name of the generated files (default: the module name)
viewsGenerate views (default true; false leaves them out)
externsHost implementations of externs (see Host interfaces)
nixPackagesPackages the externs need, resolved with nix-build by the test tooling
libsLibraries the test tooling links (-l names)

The backend option inlineRuntime copies the runtime into yaffle-runtime/ next to the output. Compile the generated source together with packages/runtime-cpp/src/*.cpp, with packages/runtime-cpp/include on the include path, as C++20 (-std=c++20). The runtime uses __int128 and GCC’s overflow builtins; it and the generated code build cleanly with -Wall -Wextra -Werror under GCC.

Generated API

#include "fmt.hpp"

fmt::Archive pf = fmt::Archive::parse(bytes);       // or parse(bytes, options)
yaffle::Bytes out = fmt::Archive::serialize(pf);       // byte-identical if unchanged
yaffle::Json j = fmt::Archive::to_json(pf);            // canonical JSON, $refs for shared values
fmt::Archive back = fmt::Archive::from_json(j);
auto o = fmt::Archive::options_from_json(args);        // endian, strict, root parameters

Every root struct X has X::Value (the struct, or std::shared_ptr<X> for shared structs), X::Options (endian, strict, and a std::optional per root parameter) and the static functions above. Errors are yaffle::Error exceptions with a code (yaffle::code_name gives "EOF", "CHECK", …), a path from the root and an offset (-1 on write).

Type mapping

yaffleC++
u8…u64, i8…i64std::uint8_t…std::uint64_t, std::int8_t…std::int64_t
u24, u40/u48/u56 (signed alike)std::uint32_t, std::uint64_t (range-checked on write)
f16/f32, f64float, double
boolbool
strings (any encoding)std::string (UTF-8)
u8[]yaffle::Bytes (std::vector<std::uint8_t>)
arraysstd::vector<T>
enumsenum class E : storage
unionsstd::variant<…> (aliases Union1, Union2, …)
structsby value; std::shared_ptr when pointers or at can target them, or on a by-value cycle
pointers to structs, arrays, unionsstd::shared_ptr (shared targets keep their identity)
pointers to scalars and stringsthe value (std::optional when nullable)
conditional and switch fields, optional inputsstd::optional

Switch cases are flattened into std::optional members (cases often share fields); from_json chooses the case by the discriminant. A one-way transform keeps its raw value in a hidden <field>_raw member, which canonical JSON shows. Structs with codec layers carry a hidden yaffle::Meta yaffle_meta_: the original encoded bytes of every codec region, reused on serialize when the decoded bytes are unchanged, so non-canonical encodings rebuild byte for byte.

Integers are computed exactly: expressions use int64_t where interval analysis proves it safe and __int128 with checked helpers otherwise. Overflow and out-of-range values are RANGE errors.

Views

auto v = fmt::Archive::view(bytes);                    // reads on access
auto es = v.entries_views();                              // child views, no reads yet
es[1].set_handler(3);                                     // same size: patched in place
es[1].set_data(bigger);                                   // structural: applied at commit
std::vector<yaffle::Patch> ps = v.patches(yaffle::Grow::Splice);
yaffle::Bytes edited = v.commit();                        // Default | Append | Splice | Error
auto whole = v.materialize();                             // every field through the view
auto fv = fmt::Archive::view_file("big.bin");          // positional reads through a block cache

A view’s struct is scanned once: scalars are read, while pointer targets, codec regions, streams and byte arrays get lazy closures, and the locations of struct values are recorded so child views need no reads. A setter patches in place when the field’s encoding keeps its size and nothing derives from it; any other edit materializes the document (one parse that records where every target was) and edits that value. commit returns the patched bytes, or re-serializes with a sticky layout that keeps targets where they were as far as the grow policy allows.

Views read through byte stores (yaffle/store.hpp): files (FileStore), decoded codec regions (decoded whole, or block by block for @ranged codecs), and from arr[*].piece streams (PieceStore: pieces located from their size fields, decoded one at a time when read).

Host interfaces

targets.cpp.externs[name] is path/file.hpp (the item has the extern’s name, with _ appended for C++ keywords), path/file.hpp#ns::item, or a bare ns::item.

  • extern fn f(T a) -> R: a callable taking the C++ types (u8[] as const yaffle::Bytes&).
  • extern codec c(args): an object with yaffle::Bytes decode(yaffle::BytesView, const yaffle::CodecCtx&) const (returning yaffle::Decoded { output, consumed } for @streaming codecs) and yaffle::Bytes encode(yaffle::BytesView, const yaffle::CodecCtx&) const. CodecCtx holds the arguments (yaffle::Arg), the decoded size when an inner @size gives it, and the block offset for @ranged codecs. async codecs may return std::future<…>; the runtime waits for it.
  • extern type X(args) : T: an object with std::optional<std::size_t> size(yaffle::BytesView, args…) const (std::nullopt: needs more bytes), T read(yaffle::BytesView, args…) const and yaffle::Bytes write(const T&, args…) const.

A host exception other than yaffle::Error becomes a CODEC error.

Runtime

Headers are in packages/runtime-cpp/include/yaffle/, one per area (yaffle.hpp includes them all), with the implementations in src/:

HeaderContents
core.hppError, paths, exact integer helpers
prim.hppinteger and float encodings
strings.hpptext encodings
json.hppthe Json value and canonical JSON conversions
codec.hppcodec calls, codec reuse (Meta), stream slices
helpers.hppsmall helpers of generated code, LEB128, stream splitting
arrays.hppderived arrays and value equality
store.hppbyte stores
reader.hppthe read context and target identity
writer.hppthe relocation-style writer: blocks, placeables, fixups, regions
view.hppdocuments, scans, view bases, materialization
harness.hppthe conformance harness (not part of yaffle.hpp)

The layout engine (src/layout.cpp) places a base region’s targets: pre-order by default, layout { … } items, pull-out claims, constraints, fixed placements and the sticky mode of view commits.

Tests

# backend tests: generate, compile with g++, run (they skip without a C++ compiler: $CXX, else g++)
YAFFLE_CPP_CACHE=/tmp/yaffle-cpp-cache npx vitest run --project yaffle packages/compiler/test/backends/cpp

# conformance, parse and view modes
YAFFLE_CPP_CACHE=/tmp/yaffle-cpp-cache \
  node --max-old-space-size=12000 --conditions=development conformance/src/cli.ts --targets cpp

runtime.test.ts runs the runtime’s own unit tests (packages/runtime-cpp/test). Builds are cached under $YAFFLE_CPP_CACHE (default <tmpdir>/yaffle-cpp-cache), keyed by content. The conformance runner compiles one program per case directory; a cold run takes a long time, and disjoint --filter groups can run in parallel processes.

Known gaps

  • No async views. Non-ranged codec regions are decoded whole when first read, and so are views over @ranged regions with layers inside the codec or several codecs.
  • Union values, and fields inside codec regions or streams, are edited structurally only. Replacing a whole array (set_entries(copy)) loses the copies’ original positions.
  • A commit after a structural edit reads the whole source and builds the new file in memory.
  • Piece-decoding errors carry the path from the struct holding the stream, not from the root.
  • Layout item arguments that depend on sizes or positions (sizeof, offsetof, $offset) are rejected.
  • A field declared in several branches with different C++ types is rejected.
  • Writing a one-way transform below the field level (an array element, a pointer target) is an INPUT error: only fields keep a raw value.
  • @origin pointers that measure later fields work on plain pointer fields only.
  • Large schemas generate a lot of C++ (about 1 MB for a dozen real-world formats), which takes minutes to compile.