Internal documentation — Work in progress
This page has not been confirmed as accurate or approved for publication. Current documentation version: 0.1.0-internal.1.
Lexical specification
Sagan source is UTF-8. Malformed UTF-8 is a lexical error. Tokens retain byte spans into the original source.
Identifiers
Textual identifiers follow Unicode 17 XID_Start and XID_Continue, with _
also accepted in either position. Emoji sequences may begin or continue an
identifier. This permits names such as:
Identifier spellings are normalized to Unicode NFC before keyword lookup and
later compiler processing. Source spans continue to reference the original
bytes. Standalone join controls and incomplete emoji ZWJ sequences are invalid.
ASCII punctuation sequences such as :) are not identifiers.
Identifiers are case-sensitive. Names beginning with __ are reserved for the
compiler. A trailing ! may be included in a method identifier as Sagan's
mutating-counterpart naming convention.
Numbers
The initial language supports decimal integers, decimal floating-point values,
scientific notation, and _ separators placed strictly between digits.
A decimal point requires a digit on both sides. Consequently .5, 5.,
.5e2, and 5.e2 are lexical errors. Binary, octal, hexadecimal, unit
suffixes, and numeric type suffixes are not currently supported.
Strings
Single and double quotes create strings. Triple double quotes create multiline
strings. Non-raw strings support escapes and ${...} interpolation; raw
strings use an r prefix and process neither feature.
The supported escapes are \\, \", \', \n, \r, \t, \0, and
\u{...}. Unicode escapes must identify a Unicode scalar value and contain no
more than six hexadecimal digits.
Comments
//begins a line comment./* ... */is a nestable block comment.///and/** ... */produce documentation-comment tokens.- Other comments are discarded by the tokenizer.
Newlines
LF and CRLF each represent one logical newline. A lone carriage return is a
lexical error. Newlines are suppressed inside parentheses and brackets and
after tokens that leave an expression incomplete. Ambiguous brace and angle
delimiter contexts are preserved for the parser to interpret.
Keywords
let fun class face enum
if else match case for in while until break continue return yield
import from as module export
hope unless finally scream
and or not self is has
true false inf nan
Literals
Decimal integers and floating-point literals support _ only between digits.
Floating-point forms may contain a decimal point with digits on both sides and
an e or E exponent with optional sign. .5, 5., malformed
separators, incomplete exponents, and letters directly following a number are
not accepted as one valid numeric literal.
Single- and double-quoted strings are recognized. Triple double quotes create
multiline strings. Non-raw strings process \\, \", \', \n,
\r, \t, \0, and \u{...}; they tokenize ${...}
interpolation, including nested braces. Raw strings begin with r and do not
process escapes or interpolation.
Comments and whitespace
// starts a line comment. Block comments /* ... */ nest.
/// and /** ... */ produce DOC_COMMENT; ordinary comments are
discarded. Spaces, tabs, form feeds, and vertical tabs are non-significant.
LF and CRLF each count as one logical newline. Blank and comment-only lines do not emit newline tokens. Newlines are suppressed inside parentheses and brackets and after a continuation token. Brace and angle-bracket newlines are preserved so the parser can interpret them with grammar context.
Punctuation and operators
Longest valid token wins. Standalone @ and backticks are errors; Schematic's
special tag, event, and code-block tokens are not part of Sagan.
Token categories
The token vocabulary also includes newline; ordinary and method identifiers; integer and float; whole/raw strings; string begin, segment, end, and interpolation boundaries; and documentation comments. Tokens carry source spans and text, with decoded numeric or string values where appropriate.