Overview¶
The document is a value¶
Every public semantic-node factory returns a frozen node. Each fluent method returns a new node rather than mutating the receiver:
from caxton import text
base = text(source="product")
titled = base.titled("Product")
assert base is not titled
Because the nodes are immutable, one factory function can produce many documents from different row sets without copying a shared graph.
Intent, not coordinates¶
The semantic model stores what you meant. It never stores cell addresses, layout decisions, execution state, caches or backend objects. Those belong to the compiler and the renderer, and are not part of the compatibility contract.
This is why table(...) has an optional anchor (declared intent) but no
"current row" (resolved state).
The three namespaces¶
Caxton keeps three value namespaces separate.
| Namespace | Built with | Evaluated by | Ends up in the artifact as |
|---|---|---|---|
| Row values | field(), path(), literal() |
Caxton | A literal value |
| Semantic columns | ref() |
Caxton | A literal value |
| Spreadsheet formulas | col(), table_ref(), sheet_ref() |
The spreadsheet | A live formula |
field() never resolves a column id and ref() never reads a row field, so the
two cannot be confused. A Python expression cannot depend on a formula-backed
column, because that column's value only exists once the file is opened.
literal() supplies a constant Python row value and never reads the current row.
Identity, source and title are separate¶
A column's id is its semantic identity, used by ref(), col(), totals,
charts and the testing views. It is independent of:
source— where the raw value comes from;title— what a human sees in the header.
Reusable ColumnSchema classes preserve this identity and expose ordinary
columns through named attributes. Their columns tuple is the canonical table
order; the schema itself never enters the semantic model.
Data is lazy and its repeatability matters¶
table() coerces its input into a DataSource once. Construction, structural
validation and semantic inspection never read a row.
A source declares one of:
REITERABLE— can be iterated again (e.g. a tuple or list);ONE_SHOT— a generator or cursor that can be consumed exactly once;UNKNOWN— repeatability cannot be determined.
A second pass over a one-shot source raises DataSourceConsumedError instead of
silently yielding nothing. Features that would need an extra pass are rejected
before writing, or documented as buffering exactly once (grouped tables and
matrices).
The pipeline¶
Public API + raw inputs
↓ DataSource coercion, without reading rows
Immutable semantic model
↓ structural validation
Requirement analysis → capabilities + workbook operation
↓ resolver checks the renderer descriptor and IR compatibility
Family compiler × selected renderer capabilities
↓
Versioned read-only spreadsheet IR
↓
Renderer → OutputSink → RenderResult
Requirement analysis does not depend on the renderer. Renderer selection
happens before the target file is opened. render() uses a memory sink.
write() normalizes a path or buffer into a sink; for paths, it replaces the
destination atomically only after the backend succeeds.