Yuku
AST
The parser produces a flat, data-oriented tree. There are no boxed structs and no pointers to chase, every read is a tagged-union switch and a slice index, and tree.deinit() frees the whole tree at once.
In JavaScript, yuku-parser returns the same tree as an ESTree / TypeScript-ESTree AST, and yuku-ast walks, builds, and checks it.
Memory layout
Three arrays in one arena hold the whole tree.
Tree
nodes every node, its payload and span stored as separate columns
extras child lists, as runs of node indices
strings text, sliced from the source or interned
Four small types carry every reference inside it.
| Type | Is |
|---|---|
NodeIndex | A node’s position in nodes, .null for an absent child |
IndexRange | A child list, a start and length in extras |
String | Text, a zero-copy slice of the source or an entry in the pool |
Span | A source range in bytes |
Reading the tree
tree.root is the program node, and four calls read everything below it.
tree.data(idx) // a node's payload, a tagged union
tree.span(idx) // its source range
tree.extra(range) // a list of child nodes
tree.string(text) // the bytes of a string
A payload is a switch away. Tags are snake_case, and ast.zig defines every one with its fields.
switch (tree.data(idx)) {
.variable_declaration => |decl| {
for (tree.extra(decl.declarators)) |declarator| {
_ = declarator;
}
},
.identifier_reference => |id| {
std.debug.print("{s}\n", .{tree.string(id.name)});
},
else => {},
}
Predicates classify a payload in one call: isExpression, isStatement, isDeclaration, isLiteral, isPattern, isCallable, isIteration, and isTypeContext. To walk the whole tree, use the traverser.
tree.source, tree.lang, and tree.source_type hold what was parsed, and isTs(), isJsx(), and isModule() ask about them.
Building and editing
Every change goes through the tree and lands in its arena.
const name = try tree.addString("total"); // interned, deduplicated
const id = try tree.addNode(.{ .identifier_reference = .{ .name = name } }, span);
const list = try tree.addExtra(&.{id}); // a child list
tree.setData(idx, data); // replace a payload
tree.setSpan(idx, span);
tree.setIdentifierName(idx, name); // any identifier kind
const text = tree.sourceSlice(start, end); // a String over the source, no copy
try tree.addDiagnostic(.{ .message = "unused", .span = span });
ast.Tree.initEmpty(allocator) starts a tree with no source, for output built from nothing. The transform traverser makes these changes during a walk.
Tokens
.tokens = true keeps every token in tree.tokens, as the parser resolved it.
for (tree.tokens) |token| {
std.debug.print("{t} {s}\n", .{ token.tag, token.text(tree.source) });
}
token.tag.isKeyword() // also isReserved, isIdentifierLike, isNumericLiteral,
// and the operator kinds
token.tag.precedence() // operator precedence, 0 when none
token.hasLineTerminatorBefore() // what ASI reads
token.isEscaped() // also hasInvalidEscape and hasLoneSurrogates
The list ends with an eof token. Every node other than program, template_element, and jsx_empty_expression starts and ends on a token boundary, so its tokens are a contiguous run.
Comments
The
commentsoption collects every comment intotree.comments, attaches each to its node, or both. An attached comment records whether it sitsbefore,after, orinsideits node, so comments move with their nodes through transforms and codegen prints them in place.