9  The words that repeat

Why did the eighth verb take minutes, when the first one took a chapter? You have met eight verbs now, and a handful of small words that traveled between them: by, where, as, all_but, descending. This chapter is about those small words, and it is the shortest useful thing in the book. The answer to the question above is that the small words never changed.

Each of them means one thing. Not one thing per verb, one thing. When you meet a verb you have never seen, the small words in it already make sense. That is most of what makes this grammar quick to get back after a year away.

9.1 by says which rows go together

You saw by on summarize, where it made the groups. It is the same word on add, where it says which rows a value is worked out within. It is the same word again on take, where it says which rows the first few are taken from.

run('
  sales
    then summarize [total] as total([revenue]) by [product]
')
product total
Doohickey 400
Gadget 720
Sprocket 240
Widget 700
sales |> summarize(total = total(revenue), by = product)
product total
Doohickey 400
Gadget 720
Sprocket 240
Widget 700
sales >> summarize(total = total(col.revenue), by = col.product)
product total
Doohickey 400
Gadget 720
Sprocket 240
Widget 700

sales then summarize [total] as total([revenue]) by [product]

On add, the same word turns a collapse into a broadcast. The aggregation is still worked out per group, and instead of one row per group, every row gets its group’s answer, which is exactly what a share needs:

run('
  sales
    then add [share] as ([revenue] / total([revenue])) by [product]
    then pick [product, revenue, share]
')
product revenue share
Widget 100 0.1428571
Gadget 120 0.1666667
Doohickey 120 0.3000000
Widget 150 0.2142857
Widget 200 0.2857143
Gadget 300 0.4166667
Doohickey 80 0.2000000
Doohickey 200 0.5000000
Widget 50 0.0714286
Gadget 240 0.3333333
Sprocket 150 0.6250000
Sprocket 90 0.3750000
Widget 75 0.1071429
Gadget 60 0.0833333
Widget 125 0.1785714
sales |>
  add(share = revenue / total(revenue), by = product) |>
  pick(product, revenue, share)
product revenue share
Widget 100 0.1428571
Gadget 120 0.1666667
Doohickey 120 0.3000000
Widget 150 0.2142857
Widget 200 0.2857143
Gadget 300 0.4166667
Doohickey 80 0.2000000
Doohickey 200 0.5000000
Widget 50 0.0714286
Gadget 240 0.3333333
Sprocket 150 0.6250000
Sprocket 90 0.3750000
Widget 75 0.1071429
Gadget 60 0.0833333
Widget 125 0.1785714
(sales
  >> add(share = col.revenue / total(col.revenue), by = col.product)
  >> pick(col.product, col.revenue, col.share))
product revenue share
Widget 100 0.1428571
Gadget 120 0.1666667
Doohickey 120 0.3
Widget 150 0.2142857
Widget 200 0.2857143
Gadget 300 0.4166667
Doohickey 80 0.2
Doohickey 200 0.5
Widget 50 0.07142857
Gadget 240 0.3333333
Sprocket 150 0.625
Sprocket 90 0.375
Widget 75 0.1071429
Gadget 60 0.08333333
Widget 125 0.1785714

Read a product’s rows and their shares sum to one: each row’s revenue over its own group’s total. In dplyr this idiom needs a grouped mutate and an ungroup, and in SQL it needs a window function. Here it is the by you already know, sitting on a different verb. And on take, the word means the same thing a third time:

run('
  sales
    then sort [revenue] descending
    then take 1 by [product]
')
date region product quantity revenue cost
2026-01-26 West Gadget 5 300 200
2026-01-09 East Widget 8 200 80
2026-03-03 East Doohickey 5 200 125
2026-06-19 East Sprocket 10 150 120
sales |> sort(descending(revenue)) |> take(1, by = product)
date region product quantity revenue cost
2026-01-26 West Gadget 5 300 200
2026-01-09 East Widget 8 200 80
2026-03-03 East Doohickey 5 200 125
2026-06-19 East Sprocket 10 150 120
sales >> sort(descending(col.revenue)) >> take(1, by = col.product)
date region product quantity revenue cost
2026-01-26 West Gadget 5 300 200
2026-01-09 East Widget 8 200 80
2026-03-03 East Doohickey 5 200 125
2026-06-19 East Sprocket 10 150 120

Three verbs, one word, one meaning. Later you will meet by on join and on widen, and it will mean the same thing there: the columns that say which rows correspond.

9.2 as gives something a name

add [name] as [value]. The name goes first, the way assignment reads, and the same shape names the result of rename and the filler in fill_missing.

run('
  sales
    then add [margin] as ([revenue] - [cost])
    then rename [earned] as [revenue]
    then pick [product, earned, margin]
')
product earned margin
Widget 100 60
Gadget 120 40
Doohickey 120 45
Widget 150 90
Widget 200 120
Gadget 300 100
Doohickey 80 30
Doohickey 200 75
Widget 50 30
Gadget 240 80
Sprocket 150 30
Sprocket 90 18
Widget 75 45
Gadget 60 20
Widget 125 75
sales |>
  add(margin = revenue - cost) |>
  rename(earned = revenue) |>
  pick(product, earned, margin)
product earned margin
Widget 100 60
Gadget 120 40
Doohickey 120 45
Widget 150 90
Widget 200 120
Gadget 300 100
Doohickey 80 30
Doohickey 200 75
Widget 50 30
Gadget 240 80
Sprocket 150 30
Sprocket 90 18
Widget 75 45
Gadget 60 20
Widget 125 75
(sales
  >> add(margin = col.revenue - col.cost)
  >> rename(earned = col.revenue)
  >> pick(col.product, col.earned, col.margin))
product earned margin
Widget 100 60
Gadget 120 40
Doohickey 120 45
Widget 150 90
Widget 200 120
Gadget 300 100
Doohickey 80 30
Doohickey 200 75
Widget 50 30
Gadget 240 80
Sprocket 150 30
Sprocket 90 18
Widget 75 45
Gadget 60 20
Widget 125 75

9.3 all_but turns a list inside out

Wrap the names you do not want. It is the same word wherever a set of columns is chosen, which so far is pick and, in the next part, lengthen.

run('
  sales
    then pick all_but [cost, region]
')
date product quantity revenue
2025-11-03 Widget 4 100
2025-11-17 Gadget 2 120
2025-12-05 Doohickey 3 120
2025-12-22 Widget 6 150
2026-01-09 Widget 8 200
2026-01-26 Gadget 5 300
2026-02-14 Doohickey 2 80
2026-03-03 Doohickey 5 200
2026-04-11 Widget 2 50
2026-05-06 Gadget 4 240
2026-06-19 Sprocket 10 150
2026-07-08 Sprocket 6 90
2026-08-24 Widget 3 75
2026-09-15 Gadget 1 60
2026-10-02 Widget 5 125
sales |> pick(all_but(cost, region))
date product quantity revenue
2025-11-03 Widget 4 100
2025-11-17 Gadget 2 120
2025-12-05 Doohickey 3 120
2025-12-22 Widget 6 150
2026-01-09 Widget 8 200
2026-01-26 Gadget 5 300
2026-02-14 Doohickey 2 80
2026-03-03 Doohickey 5 200
2026-04-11 Widget 2 50
2026-05-06 Gadget 4 240
2026-06-19 Sprocket 10 150
2026-07-08 Sprocket 6 90
2026-08-24 Widget 3 75
2026-09-15 Gadget 1 60
2026-10-02 Widget 5 125
sales >> pick(all_but(col.cost, col.region))
date product quantity revenue
2025-11-03 Widget 4 100
2025-11-17 Gadget 2 120
2025-12-05 Doohickey 3 120
2025-12-22 Widget 6 150
2026-01-09 Widget 8 200
2026-01-26 Gadget 5 300
2026-02-14 Doohickey 2 80
2026-03-03 Doohickey 5 200
2026-04-11 Widget 2 50
2026-05-06 Gadget 4 240
2026-06-19 Sprocket 10 150
2026-07-08 Sprocket 6 90
2026-08-24 Widget 3 75
2026-09-15 Gadget 1 60
2026-10-02 Widget 5 125

9.4 descending reverses an ordering

It is a modifier on a column, in a position where an ordering is being given. So it reads the same on sort and inside rank, and beyond the column it marks, it has no options of its own on either.

run('
  sales
    then add [place] as rank([revenue] descending)
    then sort [place]
    then pick [product, revenue, place]
')
product revenue place
Gadget 300 1
Gadget 240 2
Widget 200 3
Doohickey 200 3
Widget 150 5
Sprocket 150 5
Widget 125 7
Gadget 120 8
Doohickey 120 8
Widget 100 10
Sprocket 90 11
Doohickey 80 12
Widget 75 13
Gadget 60 14
Widget 50 15
sales |>
  add(place = rank(descending(revenue))) |>
  sort(place) |>
  pick(product, revenue, place)
product revenue place
Gadget 300 1
Gadget 240 2
Widget 200 3
Doohickey 200 3
Widget 150 5
Sprocket 150 5
Widget 125 7
Gadget 120 8
Doohickey 120 8
Widget 100 10
Sprocket 90 11
Doohickey 80 12
Widget 75 13
Gadget 60 14
Widget 50 15
(sales
  >> add(place = rank(descending(col.revenue)))
  >> sort(col.place)
  >> pick(col.product, col.revenue, col.place))
product revenue place
Gadget 300 1
Gadget 240 2
Widget 200 3
Doohickey 200 3
Widget 150 5
Sprocket 150 5
Widget 125 7
Gadget 120 8
Doohickey 120 8
Widget 100 10
Sprocket 90 11
Doohickey 80 12
Widget 75 13
Gadget 60 14
Widget 50 15

9.5 where asks a question

On keep the question is about a row. Later you will see it on pick, where the question is about a column, and the difference is written down rather than guessed. Inside pick you say name or kind to name what you are asking about.

run('
  sales
    then keep where ([revenue] > 150)
')
date region product quantity revenue cost
2026-01-09 East Widget 8 200 80
2026-01-26 West Gadget 5 300 200
2026-03-03 East Doohickey 5 200 125
2026-05-06 North Gadget 4 240 160
sales |> keep(revenue > 150)
date region product quantity revenue cost
2026-01-09 East Widget 8 200 80
2026-01-26 West Gadget 5 300 200
2026-03-03 East Doohickey 5 200 125
2026-05-06 North Gadget 4 240 160
sales >> keep(col.revenue > 150)
date region product quantity revenue cost
2026-01-09 East Widget 8 200 80
2026-01-26 West Gadget 5 300 200
2026-03-03 East Doohickey 5 200 125
2026-05-06 North Gadget 4 240 160

9.6 Where by does not go

A word that travels still has a boundary, and the boundary is stated rather than guessed. by means per group, and sort already orders every row, so the word has nothing to add there and is refused by name. The sentence below is the text form, so the call is the same string in both tabs:

run('sales then sort by [region]')
Error:
! 
illegal: `sort` does not take the word `by`. Write `sort [column]`, and `descending` after it to run the other way
  |
1 | sales then sort by [region]
  |                 ^^
try:
    run('sales then sort by [region]')
except GodError as refusal:
    print(refusal)

illegal: `sort` does not take the word `by`. Write `sort [column]`, and `descending` after it to run the other way
  |
1 | sales then sort by [region]
  |                 ^^

The message says what sort does take, which is the whole of it.

9.7 Why this matters more than it looks

A vocabulary is easy to learn and easy to lose. What survives a year away is not the word list. It is the rules for putting words together, and those only survive if there are few of them and they have no exceptions.

So the test this grammar sets itself is not whether it has a word for everything. It is whether a word you learned in one place still means the same thing in the next place you meet it. Every part after this one is that test being run. And you now hold the working form of the preface’s motto. If you can say something two ways in this grammar, one of them is a bug, and the author wants to hear about it. That sentence is not rhetoric. It is the standard each of the last seven chapters answered to, from the missing ascending to the missing drop. It is falsifiable by any reader with an afternoon.

The eight verbs and five small words you now have are the everyday core of the grammar, and the preface’s promise has reached its first checkpoint. Day one was reading, and you have been reading since the first pipeline. Day two is writing, and the practice pair ahead closes with it. Those two chapters are where this part’s claims stop being the book’s and start being yours.