Yuku

AST

The parser produces a flat, data-oriented tree. There are no boxed structs and no pointers to chase, every read is a tagged-union switch and a slice index, and tree.deinit() frees the whole tree at once.

In JavaScript, yuku-parser returns the same tree as an ESTree / TypeScript-ESTree AST, and yuku-ast walks, builds, and checks it.

Memory layout

Three arrays in one arena hold the whole tree.

Tree
 nodes    every node, its payload and span stored as separate columns
 extras   child lists, as runs of node indices
 strings  text, sliced from the source or interned

Four small types carry every reference inside it.

TypeIs
NodeIndexA node’s position in nodes, .null for an absent child
IndexRangeA child list, a start and length in extras
StringText, a zero-copy slice of the source or an entry in the pool
SpanA source range in bytes

Reading the tree

tree.root is the program node, and four calls read everything below it.

tree.data(idx)    // a node's payload, a tagged union
tree.span(idx)    // its source range
tree.extra(range) // a list of child nodes
tree.string(text) // the bytes of a string

A payload is a switch away. Tags are snake_case, and ast.zig defines every one with its fields.

switch (tree.data(idx)) {
    .variable_declaration => |decl| {
        for (tree.extra(decl.declarators)) |declarator| {
            _ = declarator;
        }
    },
    .identifier_reference => |id| {
        std.debug.print("{s}\n", .{tree.string(id.name)});
    },
    else => {},
}

Predicates classify a payload in one call: isExpression, isStatement, isDeclaration, isLiteral, isPattern, isCallable, isIteration, and isTypeContext. To walk the whole tree, use the traverser.

tree.source, tree.lang, and tree.source_type hold what was parsed, and isTs(), isJsx(), and isModule() ask about them.

Building and editing

Every change goes through the tree and lands in its arena.

const name = try tree.addString("total");          // interned, deduplicated
const id = try tree.addNode(.{ .identifier_reference = .{ .name = name } }, span);
const list = try tree.addExtra(&.{id});            // a child list
tree.setData(idx, data);                           // replace a payload
tree.setSpan(idx, span);
tree.setIdentifierName(idx, name);                 // any identifier kind
const text = tree.sourceSlice(start, end);         // a String over the source, no copy
try tree.addDiagnostic(.{ .message = "unused", .span = span });

ast.Tree.initEmpty(allocator) starts a tree with no source, for output built from nothing. The transform traverser makes these changes during a walk.

Comments

The comments option collects every comment into tree.comments, attaches each to its node, or both. An attached comment records whether it sits before, after, or inside its node, so comments move with their nodes through transforms and codegen prints them in place.

for (tree.commentsOf(idx)) |c| {
    std.debug.print("{t} {s}\n", .{ c.position, tree.string(c.value) });
}

Tokens

.tokens = true keeps every token in tree.tokens, as the parser resolved it.

for (tree.tokens) |token| {
    std.debug.print("{t} {s}\n", .{ token.tag, token.text(tree.source) });
}
token.tag.isKeyword()             // also isReserved, isIdentifierLike, isNumericLiteral,
                                  // and the operator kinds
token.tag.precedence()            // operator precedence, 0 when none
token.hasLineTerminatorBefore()   // what ASI reads
token.isEscaped()                 // also hasInvalidEscape and hasLoneSurrogates

The list ends with an eof token. Every node other than program, template_element, and jsx_empty_expression starts and ends on a token boundary, so its tokens are a contiguous run.