orders = pd.DataFrame({"order date": ["2026-01-02", "2026-01-05"],
"total": [40, 90]})
orders >> keep(col["order date"] > "2026-01-03")| order date | total |
|---|---|
| 2026-01-05 | 90 |
A pipeline is the same sentence in both languages. Your first pipeline showed the two places they differ. This appendix is the rest of it: why the pipe is what it is, the three rules Python adds, and how R’s own habits are read. Then it turns to names: the ones god takes over, the ones other packages take back, and the one god shares with gog.
Everything else in the book is written once, with a tab for each language, because everywhere else the two really are the same sentence.
>> is the pipe in Python, and it is an operator rather than a method call for the reason the preface gives. An arrow points at the next step, and a dot ends a sentence. Python cannot have |> itself. Its operator set is fixed by the language, and |> does not even tokenize.
Only two operators are left free by pandas, polars and PySpark alike: >> and <<, the same arrow in two directions. The one that points at the next step is the pipe, which makes it the choice rather than a preference. | was considered and rejected, because pandas defines it, so frame | verb never reaches the verb at all. It tries an elementwise or with the verb as one more value, and the error that stops it comes from inside pandas, saying nothing about the pipe you meant.
col.name is how Python names a column. R hands a verb the expression you wrote, unevaluated, so a bare revenue can be read as a column. Python evaluates every argument before god sees it: a bare revenue there is a session variable or an error, so a column says that it is one.
A name that is not a Python identifier is still a column. Write it in brackets, col["order date"], and it is the same expression the dot would have made:
orders = pd.DataFrame({"order date": ["2026-01-02", "2026-01-05"],
"total": [40, 90]})
orders >> keep(col["order date"] > "2026-01-03")| order date | total |
|---|---|
| 2026-01-05 | 90 |
R’s spelling for the same name is its own, backticks, and R removes them before god ever sees the name:
orders <- data.frame(`order date` = c("2026-01-02", "2026-01-05"),
total = c(40, 90), check.names = FALSE)
orders |> keep(`order date` > "2026-01-03")| order date | total |
|---|---|
| 2026-01-05 | 90 |
None of them comes from the grammar, and this book states them here rather than letting you find them in an error message.
Comparisons joined by & need parentheses, because & binds more tightly than a comparison in Python. Negation is ~, because not is a keyword whose result Python turns into a bool, so an object cannot see it.
sales >> keep((col.revenue > 150) & (col.cost < 100))| date | region | product | quantity | revenue | cost |
|---|---|---|---|---|---|
| 2026-01-09 | East | Widget | 8 | 200 | 80 |
sales >> keep(~(col.product == "Widget"))| date | region | product | quantity | revenue | cost |
|---|---|---|---|---|---|
| 2025-11-17 | East | Gadget | 2 | 120 | 80 |
| 2025-12-05 | West | Doohickey | 3 | 120 | 75 |
| 2026-01-26 | West | Gadget | 5 | 300 | 200 |
| 2026-02-14 | North | Doohickey | 2 | 80 | 50 |
| 2026-03-03 | East | Doohickey | 5 | 200 | 125 |
| 2026-05-06 | North | Gadget | 4 | 240 | 160 |
| 2026-06-19 | East | Sprocket | 10 | 150 | 120 |
| 2026-07-08 | West | Sprocket | 6 | 90 | 72 |
| 2026-09-15 | East | Gadget | 1 | 60 | 40 |
Asking whether a value is one of several is a method, since Python’s in cannot be reached either.
sales >> keep(col.region.is_in(["West", "East"]))| date | region | product | quantity | revenue | cost |
|---|---|---|---|---|---|
| 2025-11-03 | West | Widget | 4 | 100 | 40 |
| 2025-11-17 | East | Gadget | 2 | 120 | 80 |
| 2025-12-05 | West | Doohickey | 3 | 120 | 75 |
| 2026-01-09 | East | Widget | 8 | 200 | 80 |
| 2026-01-26 | West | Gadget | 5 | 300 | 200 |
| 2026-03-03 | East | Doohickey | 5 | 200 | 125 |
| 2026-04-11 | West | Widget | 2 | 50 | 20 |
| 2026-06-19 | East | Sprocket | 10 | 150 | 120 |
| 2026-07-08 | West | Sprocket | 6 | 90 | 72 |
| 2026-09-15 | East | Gadget | 1 | 60 | 40 |
| 2026-10-02 | West | Widget | 5 | 125 | 50 |
Write the values as a list when the order matters to you. A set works too, and gets sorted before it is written, because Python’s hashing changes between runs and the same pipeline has to read the same way every time.
Asking whether a value is missing is a method of the same shape: col.cost.is_missing(), where R writes is.na(cost).
The third rule is a refusal rather than a spelling. Python’s own and, or, in and if all ask a value to become one plain yes or no, and a column expression refuses, because its honest answer is one per row:
try:
if col.revenue > 150:
print("large")
except GodError as refusal:
print(refusal)this is a column expression, not a yes or no. Combine conditions with `&`, `|` and `~`, and hand the whole expression to a verb
The message names &, | and ~ at the moment you reach the wrong spelling, which is earlier than a wrong answer would have told you.
One word used to be on this list and no longer is. The marker that inverts a pick was except, which Python had to spell except_, because except is a keyword here. It is all_but in both languages now, and in the text form. No word in the vocabulary is spelled differently in the two languages.
Two host-side calls do differ, and neither is a word of the vocabulary. Asking a pipeline for its text form is format(p) in R and p.written() in Python. format is the word R already owns for turning a thing into text, and Python’s format means something else, so its method says what it gives. And show_as prints in R while it returns in Python, for the notebook reason What it wrote explains.
The grammar has one spelling for each idea, and some of R’s spellings are not it. You still write R, and the verbs translate as they build the sentence. format shows what they wrote.
cat(format(sales |> keep(region != "West" & !is.na(cost))))sales
then keep where (([region] is not "West") and ([cost] is not missing))
!= became is not and is.na became is missing. The everyday replacements fit in one table.
| You write | god writes |
|---|---|
| == | is |
| != | is not |
| & | and |
| | | or |
| ! | not |
| %in% | in { } |
| is.na(x) | x is missing |
| TRUE | yes |
| NA | missing |
R’s text spellings translate the same way: startsWith becomes starts, endsWith becomes ends, grepl becomes contains, and tolower and toupper become lower and upper.
A set of values is written out, because the grammar has no variables yet and will not guess at one.
sales |>
keep(region %in% c("West", "East")) |>
summarize(n = row_count(), by = region)| region | n |
|---|---|
| East | 5 |
| West | 6 |
god takes overAttaching god in R replaces sort for the rest of your session. That is deliberate: a prefix on every sentence is a worse trade than a good message on the few calls that go wrong. The message names what was shadowed and how to reach it.
sort(c(3, 1, 2))Error:
! god's `sort` orders the rows of a table, and `c(3, 1, 2)` is not a table.
For R's own, write `base::sort(c(3, 1, 2))`.
What it does not take over is every other package you have. R looks a name up inside a package through that package’s own namespace, never through the search path you attached to, so a function that calls sort in its own body is untouched. median and quantile both do.
c(median(c(3, 1, 2)), quantile(1:9, 0.5)) 50%
2 5
So the cost is your own calls to sort, at the console or in a script, and nothing further. base::sort is always there, and the message above says so at the moment you need it.
Python has no attach step. import god adds one name to your session, so it collides with nothing at all, and only from god import * can.
Four of god’s names are dplyr’s as well: collect, pick, rename and summarize. keep is purrr’s. Attach one of those packages after god and its version of the name is what your next line reaches.
Three of the four cost you nothing. collect, rename and summarize are generics, and god gives each of them a method for a pipeline, so the sentence runs whichever package you attached last.
sales |>
keep(region == "West") |>
dplyr::summarize(takings = total(revenue), by = product)| product | takings |
|---|---|
| Doohickey | 120 |
| Gadget | 300 |
| Sprocket | 90 |
| Widget | 275 |
That is dplyr’s summarize, handed a god pipeline and answering as god’s. Which order you wrote your library() calls in is not visible from here.
pick is the exception, and the reason belongs to dplyr rather than to god: its pick is not a generic, so there is no method to give it. Write god::pick when both packages are attached.
keep after purrr is the same case: purrr’s keep is not a generic either, so attaching purrr last takes the word back. The repair is the same shape, god::keep.
gogcol is a good enough name that other people picked it too. polars uses it, and so does gog, this project’s sibling for drawing plots, where it does the same job for the same reason. from god import * and from gog import * in one file means the second one wins and the first is gone.
Import them by name when you want both, which is what Python’s namespaces are for:
import god
import gog
sales_by_product = god.collect(
sales >> god.summarize(revenue = god.total(god.col.revenue), by = god.col.product)
)Or bring in one of them and reach for the other by name. The failure is loud if you forget: god’s verbs cannot read a gog column, so a mixed-up pipeline stops rather than guessing.
god you haveEach language answers in its own idiom, and a release is refused when the two would disagree.
packageVersion("god")[1] '0.2.2'
import god
god.__version__'0.2.2'
Same number, two idioms, which is this appendix in one line: the differences between the languages are spellings, and never the sentence.