show_as(sales |> keep(region == "West") |> take(10), "dplyr")sales |>
filter((region == "West")) |>
head(10)
A small vocabulary covers most of what people do and never all of it, so the question is not whether you reach its edge but what happens there. Ask any pipeline what it would be in dplyr, or ask for the query. There are seven answers: sql, spark, dplyr, pandas, polars, pyspark, and god itself.
show_as(sales |> keep(region == "West") |> take(10), "dplyr")sales |>
filter((region == "West")) |>
head(10)
show_as(sales >> keep(col.region == "West") >> take(10), "dplyr")sales |>
filter((region == "West")) |>
head(10)
show_as also takes the sentence itself, with any tables passed by name, as show_as('sales then take 3', "dplyr", sales = sales); the phrasebook appendix is generated that way. And the two languages hand the result back differently, on purpose. R prints the code and returns it invisibly, the way R’s print methods behave. Python returns it as a string, for print or for storing, because printing and returning both made a notebook say everything twice.
show_as(sales |> summarize(revenue = total(revenue), by = product), "sql")WITH step0 AS (SELECT * FROM "sales"),
step1 AS (SELECT "product", sum("revenue") AS "revenue" FROM step0 GROUP BY "product" ORDER BY "product" NULLS LAST)
SELECT * FROM step1
show_as(sales >> summarize(revenue = total(col.revenue), by = col.product),
"sql")WITH step0 AS (SELECT * FROM "sales"),
step1 AS (SELECT "product", sum("revenue") AS "revenue" FROM step0 GROUP BY "product" ORDER BY "product" NULLS LAST)
SELECT * FROM step1
The query has a shape nobody would write by hand, and the shape is deliberate. Each then becomes one named step in a WITH chain, so the sentence and the query match line for line, and a longer pipeline is more named steps rather than a deeper nest. A person would fuse steps and name them after their meaning; a machine’s first duty is that you can check its work. The engine flattens the chain into one plan before reading a row, so the shape costs nothing at run time.
"sql" is the query for the engine on your machine. "spark" is the same pipeline for Spark, and the point is what does not change: the sentence you wrote.
show_as(sales |> summarize(revenue = total(revenue), by = product), "spark")WITH step0 AS (SELECT * FROM `sales`),
step1 AS (SELECT `product`, sum(`revenue`) AS `revenue` FROM step0 GROUP BY `product` ORDER BY `product` NULLS LAST)
SELECT * FROM step1
show_as(sales >> summarize(revenue = total(col.revenue), by = col.product),
"spark")WITH step0 AS (SELECT * FROM `sales`),
step1 AS (SELECT `product`, sum(`revenue`) AS `revenue` FROM step0 GROUP BY `product` ORDER BY `product` NULLS LAST)
SELECT * FROM step1
The two dialects differ in nine ways, and the queries above show the first. A column is `revenue` on Spark and "revenue" here, and that one matters. A double-quoted name on Spark is not a column at all: it is the text 'revenue'. A query written for the wrong engine would therefore run without error and put the column’s own name in every row. Two more are spelling: dropping a column is EXCEPT rather than EXCLUDE, and stopping a query is raise_error rather than error. A backslash inside a text value has to be written twice on Spark. The word weekday is asked for through a different function on each engine, so that Monday can stay 1 on both. latest, a rolling median, and join_rows are three more, each written differently on the two engines. And a widen that names no columns is a sentence Spark cannot be given at all, which is the refusal staged later on this page.
None of that is yours to remember. It is the difference between writing a pipeline and writing a query.
A query is one answer. The other is the dataframe library you already use, which is what you want when the table is in front of you and there is no engine to hand anything to. The same pipeline, three ways:
show_as(sales |>
add(margin = revenue - cost) |>
keep(margin > 50) |>
sort(descending(margin)), "polars")(sales
.with_columns((pl.col("revenue") - pl.col("cost")).alias("margin"))
.filter((pl.col("margin") > 50))
.sort(["margin"], descending=[True], nulls_last=True))
show_as(sales
>> add(margin = col.revenue - col.cost)
>> keep(col.margin > 50)
>> sort(descending(col.margin)), "polars")(sales
.with_columns((pl.col("revenue") - pl.col("cost")).alias("margin"))
.filter((pl.col("margin") > 50))
.sort(["margin"], descending=[True], nulls_last=True))
show_as(sales |>
add(margin = revenue - cost) |>
keep(margin > 50) |>
sort(descending(margin)), "pandas")(sales
.assign(margin=lambda d: (d["revenue"] - d["cost"]))
.loc[lambda d: (d["margin"] > 50)]
.sort_values(["margin"], ascending=[False]))
show_as(sales
>> add(margin = col.revenue - col.cost)
>> keep(col.margin > 50)
>> sort(descending(col.margin)), "pandas")(sales
.assign(margin=lambda d: (d["revenue"] - d["cost"]))
.loc[lambda d: (d["margin"] > 50)]
.sort_values(["margin"], ascending=[False]))
show_as(sales |>
add(margin = revenue - cost) |>
keep(margin > 50) |>
sort(descending(margin)), "pyspark")(sales
.withColumn("margin", (F.col("revenue") - F.col("cost")))
.filter((F.col("margin") > 50))
.orderBy(F.col("margin").desc_nulls_last()))
show_as(sales
>> add(margin = col.revenue - col.cost)
>> keep(col.margin > 50)
>> sort(descending(col.margin)), "pyspark")(sales
.withColumn("margin", (F.col("revenue") - F.col("cost")))
.filter((F.col("margin") > 50))
.orderBy(F.col("margin").desc_nulls_last()))
Each verb became one method, and the order of the steps did not move. What changes is how a column is named. polars writes pl.col, PySpark writes F.col, and pandas has no name of its own, so it borrows the frame it was handed and writes d["margin"] inside a small function.
That last one is why the pandas rendering reads the least like the other two, and it is the honest answer rather than a failure to find a plainer one. pandas has no single way to chain steps, so the idiom that does chain is .assign and .loc[lambda d: ...], and a column inside either of those has to name the frame again. The grammar says [margin] once.
One R dialect is missing from the list on purpose, and the reason is meaning rather than effort. Idiomatic data.table works by reference: := changes the table in place, so pasted code would quietly modify your data, which no god pipeline may ever do. A faithful translation would have to open with copy() and stop reading like the data.table its users write, and a translation neither side recognizes serves nobody.
Where an engine cannot say what a sentence means, you get a refusal rather than a query that says something close. Spark has to be told which columns a widen makes, because its pivot cannot find them in the data:
show_as(marks |> widen(name = question, value = mark, by = student),
"spark")Error:
!
illegal: Spark has to be told which columns a `widen` makes, and this one takes them from the data. Say what it makes: `giving [q1, q2, q3]`
|
2 | then widen name [question], value [mark] by [student]
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
try:
show_as(marks
>> widen(name = col.question, value = col.mark, by = col.student),
"spark")
except GodError as refusal:
print(refusal)
illegal: Spark has to be told which columns a `widen` makes, and this one takes them from the data. Say what it makes: `giving [q1, q2, q3]`
|
2 | then widen name [question], value [mark] by [student]
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
show_as starts from a pipeline and a table you are holding. god_sql is the shortcut for the other situation: you have the sentence and you know the columns. The query itself is the thing you are after, for a scheduler, a dashboard, or a warehouse console that will run it far from here.
cat(god_sql(
'sales then keep where [region] is "West" then take 3',
"region:text,product:text,revenue:number"
))WITH step0 AS (SELECT * FROM "sales"),
step1 AS (SELECT * FROM step0 WHERE ("region" = 'West')),
step2 AS (SELECT * FROM step1 LIMIT 3)
SELECT * FROM step2
print(god_sql(
'sales then keep where [region] is "West" then take 3',
"region:text,product:text,revenue:number"
))WITH step0 AS (SELECT * FROM "sales"),
step1 AS (SELECT * FROM step0 WHERE ("region" = 'West')),
step2 AS (SELECT * FROM step1 LIMIT 3)
SELECT * FROM step2
No table is attached and none is needed. The column list is the contract, and the same checks run against it before anything is written out.
Everything on this page is one binary, god-cli, which the transcripts in this book shorten to god. Both installed packages carry a copy, and the repository builds one with cargo build --release; put it on your PATH and the transcripts run as shown. It reads a sentence and writes text, so it needs no R and no Python: give it the columns, and name the target with --as.
$ god --columns 'region:text,revenue:number' --as dplyr 'sales then keep where [region] is "West" then take 3'
sales |>
filter((region == "West")) |>
head(3)
Wrap the pipeline in single quotes, since a text value inside it uses double ones. The pipeline can also arrive on standard input, as god [options] < pipeline.god, which is the shape a scheduled job wants.
Two smaller questions have flags of their own. --needs prints the tables a pipeline reads and stops, which is how a launcher knows what to describe before describing anything:
$ god --needs 'sales then join products by [product] then take 3'
sales
products
And --vocabulary prints every word the grammar has, one per line, tagged with its kind; Appendix B is generated from it.
The exit code says what happened, so a script can ask. 0 is an answer, on standard output. 1 is a command line the tool could not read, a flag given twice included, with the usage on standard error. 2 is a refusal:
$ god --columns 'region:text' --as excel 'sales then take 1'
illegal: there is no backend called `excel`. There is: sql, spark, dplyr, pandas, polars, pyspark, god
An assumption is a note rather than a refusal, and it costs nothing. Where the grammar decided something you did not say, the note lands on standard error while the answer still arrives with exit 0, so nothing is decided silently.
god_sql writes the query and stops. run writes it, executes it, and gives back the table, starting from a sentence instead of from a pipeline you built with verbs.
run('sales then keep where [region] is "West" then take 3')| date | region | product | quantity | revenue | cost |
|---|---|---|---|---|---|
| 2025-11-03 | West | Widget | 4 | 100 | 40 |
| 2025-12-05 | West | Doohickey | 3 | 120 | 75 |
| 2026-01-26 | West | Gadget | 5 | 300 | 200 |
run('sales then keep where [region] is "West" then take 3') date region product quantity revenue cost
0 2025-11-03 West Widget 4 100 40
1 2025-12-05 West Doohickey 3 120 75
2 2026-01-26 West Gadget 5 300 200
Read what sits inside the parentheses in each tab. It is the same string, to the character. Everywhere else in this book the two languages differ in the pipe glyph and in how a column is named. A sentence in the text form has neither of those, so nothing is left to respell.
The table at the head of the sentence is looked up where you are calling from, which is how sales was found here. Pass it by name when it is somewhere else, as run(pipeline, sales = last_quarter).
The grammar has exactly one rule about quotes: a text value inside a sentence is written with double quotes, always. Every other quote you will ever type belongs to your own language, and your language’s habits apply. Three seats, one rule each:
'West' and "West" as the same string, and that stays true here. The grammar receives the value, never the marks around it.[region] is "West". The sentence is the grammar’s, and the grammar has one spelling.Look at the first seat, because the rule about double quotes is easy to over-apply. In the native spellings, your language’s quotes are your language’s business:
sales |> keep(region == 'West') |> take(2)| date | region | product | quantity | revenue | cost |
|---|---|---|---|---|---|
| 2025-11-03 | West | Widget | 4 | 100 | 40 |
| 2025-12-05 | West | Doohickey | 3 | 120 | 75 |
sales >> keep(col.region == 'West') >> take(2)| date | region | product | quantity | revenue | cost |
|---|---|---|---|---|---|
| 2025-11-03 | West | Widget | 4 | 100 | 40 |
| 2025-12-05 | West | Doohickey | 3 | 120 | 75 |
Both tabs answer, single quotes and all, because by the time the grammar sees this sentence the value is "West" either way. This book writes double quotes in host code as one habit fewer, and nothing depends on it.
Around a sentence, single quotes are the transcripts’ convention, kept for the same reason everywhere. The sentence’s own double quotes need no escape, and the string pastes into a shell, an R script or a Python one unchanged.
Swap the two out of habit and the grammar catches the half that reaches it. A text value in single quotes is refused with the spelling named, so the mistake costs a message rather than a wrong answer:
run("sales then keep where [region] is 'West' then take 3")Error:
!
illegal: the grammar writes a text value with double quotes, and only double quotes, so there is one spelling rather than two. Write `"West"`
|
1 | sales then keep where [region] is 'West' then take 3
| ^
try:
run("sales then keep where [region] is 'West' then take 3")
except GodError as refusal:
print(refusal)
illegal: the grammar writes a text value with double quotes, and only double quotes, so there is one spelling rather than two. Write `"West"`
|
1 | sales then keep where [region] is 'West' then take 3
| ^
The other half belongs to the hosts: a single quote inside a single-quoted string is a syntax error R and Python report themselves, before god is ever called. And when a text value genuinely needs an apostrophe, each host has its own answer. Python’s triple quotes, run('''…'''), and R’s raw string, run(r"(…)"), both hold either kind of quote unescaped.
A sentence is data. It fits in a database cell, a configuration file, a spreadsheet column, or a message to a colleague. So the person who writes a question and the person who runs it do not have to be the same person.
Someone whose work is reading tables, and who will never write R or Python, can still write this one and say whether it asks the right thing:
sales
then keep where [region] is "West" # this quarter's question
then summarize [total] as total([revenue]) by [product]
The # starts a comment that runs to the end of its line, here exactly as it does in both languages, so a stored sentence can carry its reasoning along. An analyst calls run on it later, or a scheduled job does. The sentence crosses that handoff intact, because it is not code in either language, and it is still readable by the person who wrote it.
One boundary matters here. run is an R function and a Python function, so calling it needs one of those two languages installed. What travels without them is the sentence, not the call around it. The next chapter follows the sentence to tables that are not in front of you.