This appendix answers the questions a technical reader asks of any table tool. Where does the work happen, when does it happen, who optimizes it, and what happens when the data outgrows the machine? The short answer to all four is one sentence: god compiles, and an engine executes.
H.1 One engine does the work
A pipeline here is compiled into a single query, a few milliseconds of work, and the appendix on speed clocks it live. The query is then executed by software built for executing queries: DuckDB on this machine, Spark SQL or a warehouse wherever Somewhere else points. god never touches a row. It owns the sentence, the checking, and the writing out, while the scanning, joining and grouping are the engine’s. That is why the appendix on speed credits the engine by name.
The whole pipeline reaches the engine as one query, never a step at a time:
WITH step0 AS (SELECT * FROM "sales"),
step1 AS (SELECT * FROM step0 WHERE ("region" = 'West')),
step2 AS (SELECT *, ("revenue" - "cost") AS "margin" FROM step1),
step3 AS (SELECT "product", sum("margin") AS "margin" FROM step2 GROUP BY "product" ORDER BY "product" NULLS LAST),
step4 AS (SELECT * FROM step3 ORDER BY "margin" DESC NULLS LAST),
step5 AS (SELECT * FROM step4 LIMIT 10)
SELECT * FROM step5
WITH step0 AS (SELECT * FROM "sales"),
step1 AS (SELECT * FROM step0 WHERE ("region" = 'West')),
step2 AS (SELECT *, ("revenue" - "cost") AS "margin" FROM step1),
step3 AS (SELECT "product", sum("margin") AS "margin" FROM step2 GROUP BY "product" ORDER BY "product" NULLS LAST),
step4 AS (SELECT * FROM step3 ORDER BY "margin" DESC NULLS LAST),
step5 AS (SELECT * FROM step4 LIMIT 10)
SELECT * FROM step5
Five steps, one text. The engine sees the whole question before it reads a single row.
H.2 Always lazy, one surface
Some tools keep two interfaces: an eager one that runs each step as it is typed, and a lazy one that plans. god has one, and it is the lazy one. A verb returns a plan, printing runs it, and collect hands over the table, which is the whole of Nothing runs until you ask. There is no eager mode to be caught in by accident, and the difference is not cosmetic: an engine given steps one at a time cannot plan across them. The dated record in the appendix on speed carries an eager and a lazy spelling of the same library, and the gap between them is the price of running step by step.
H.3 What the engine optimizes
Because the query arrives whole, the planning belongs to the engine: which filter to push into which scan, which join to run first, which columns never to read at all. Those are decades of database work, and god’s job is to stay out of their way. The steps it writes are flattened by the engine into one plan, so a pipeline of eight steps is not eight passes over the data. Nothing about this needed a word in the vocabulary, which is the point: the grammar buys the optimizer by handing over text.
H.4 When the data outgrows the machine
Two sizes of big, two answers.
A table bigger than memory, held as a frame: the engine can spill its work to disk. But a pipeline here begins at a table the host already holds, so on one machine the frame is as large as the data can get. god has no word for reading a file into a pipeline, and the coverage appendix records that gap beside the others, with what to do today.
A table bigger than the machine is the second answer, and it is the warehouse and the cluster named at the top of this page. The table stays where it is, the sentence goes there as text, and no data moves at all. At that size the honest question is not whether one process can stream one file, and the grammar’s answer is to run where the data lives.
Compiling to text has one more property: everything above is inspectable. show_as prints the exact query the engine will run, and the same pipeline in every other target beside it, so nothing on this page is taken on faith.