1  Your first pipeline

You have a year of sales in front of you. Which product earned the most in the West, after costs? That is a real question, one of a hundred like it you will ask this year, and one sentence answers it.

A pipeline is that sentence. It names a table, then says what happens to it, one step at a time, in the order the steps happen. Here is the whole answer. Read it aloud before you read anything about it.

run('
  sales
    then keep where ([region] is "West")
    then add [margin] as ([revenue] - [cost])
    then summarize [margin] as total([margin]),
         [orders] as row_count() by [product]
    then sort [margin] descending
')
product margin orders
Widget 165 3
Gadget 100 1
Doohickey 45 1
Sprocket 18 1
sales |>
  keep(region == "West") |>
  add(margin = revenue - cost) |>
  summarize(margin = total(margin), orders = row_count(), by = product) |>
  sort(descending(margin))
product margin orders
Widget 165 3
Gadget 100 1
Doohickey 45 1
Sprocket 18 1
(sales
  >> keep(col.region == "West")
  >> add(margin = col.revenue - col.cost)
  >> summarize(margin = total(col.margin), orders = row_count(),
               by = col.product)
  >> sort(descending(col.margin)))
product margin orders
Widget 165 3
Gadget 100 1
Doohickey 45 1
Sprocket 18 1

sales then keep where [region] is "West" then add [margin] as [revenue] - [cost] then summarize [margin] as total([margin]), [orders] as row_count() by [product] then sort [margin] descending

Keep the West rows. Add a margin. Total the margin by product and count the orders. Sort by margin, largest first. That is the whole sentence. The table under it is what the engine returned when this page was built, and its top row is the answer to the question you started with.

Notice what you did not have to know. No method names, no order of clauses to memorize. You read a sentence in plain English, and if you can read this one you can already read most of the pipelines in this book, including the ones no page has shown you yet.

1.1 The gloss, and why it runs

The italic line under the example is the same sentence again, in the grammar’s own words. This book prints one under the first example of every teaching chapter. Reading it aloud is not a ritual for its own sake: it is the preface’s retention promise being rehearsed. A sentence you can say is a sentence you can still say after a year away.

The gloss is not decoration, and it is not pseudocode. The text form is the grammar’s main spelling, the one run takes, and the one this book points a new reader at first. It is also the only spelling that still means something where there is neither R nor Python: a database cell, a configuration file, a stored pipeline. In it, then is the pipe, written as a word. A column goes in square brackets, [margin], so that any name at all can be a column. And because it is real, this book’s build hands every gloss on every page to the engine and refuses to publish if one stops parsing. A gloss cannot quietly go stale; the build will not let it.

1.2 The table

Every example in this book runs on a table called sales, unless a chapter says otherwise. It is fifteen rows: three regions, four products, a year of dates, and for each order a quantity, the revenue, and the cost. That is small enough to check any total by hand, and you should. The whole cast of tables this book uses is declared in the preface, along with how to put the same rows in your session. None of them changes shape from one chapter to the next.

date region product quantity revenue cost
2025-11-03 West Widget 4 100 40
2025-11-17 East Gadget 2 120 80
2025-12-05 West Doohickey 3 120 75
2025-12-22 North Widget 6 150 60
2026-01-09 East Widget 8 200 80
2026-01-26 West Gadget 5 300 200
2026-02-14 North Doohickey 2 80 50
2026-03-03 East Doohickey 5 200 125
2026-04-11 West Widget 2 50 20
2026-05-06 North Gadget 4 240 160
2026-06-19 East Sprocket 10 150 120
2026-07-08 West Sprocket 6 90 72
2026-08-24 North Widget 3 75 30
2026-09-15 East Gadget 1 60 40
2026-10-02 West Widget 5 125 50
date region product quantity revenue cost
2025-11-03 West Widget 4 100 40
2025-11-17 East Gadget 2 120 80
2025-12-05 West Doohickey 3 120 75
2025-12-22 North Widget 6 150 60
2026-01-09 East Widget 8 200 80
2026-01-26 West Gadget 5 300 200
2026-02-14 North Doohickey 2 80 50
2026-03-03 East Doohickey 5 200 125
2026-04-11 West Widget 2 50 20
2026-05-06 North Gadget 4 240 160
2026-06-19 East Sprocket 10 150 120
2026-07-08 West Sprocket 6 90 72
2026-08-24 North Widget 3 75 30
2026-09-15 East Gadget 1 60 40
2026-10-02 West Widget 5 125 50

1.3 The same sentence, twice

The R and Python tabs of the first example are not translations of each other. They are the same sentence, and only two things about them differ.

The first is the pipe. R writes |>, which is R’s own. Python writes >>, because Python cannot have |> at all, and of the operators a data frame leaves free that is the one available everywhere.

The second is how a column is named. R can look a name up in your data before it looks in your session, so a bare revenue is unambiguous there. Python has no such hook, so a column says that it is one: col.revenue.

Every verb, every keyword and every argument position is the same. Someone who learned the grammar in one language has learned it in the other. Every tab on every page of this book is executed when the book is built, so that claim is checked on every page rather than asserted once here.

Appendix A has the rest of the differences between the two languages, and you do not need it yet.

1.4 What a step is

Each step in the sentence takes a table and gives a table back. That one fact matters more than it looks like it should. It means you can stop the pipeline anywhere and look at what you have so far. It means the order you read is the order things happen, top to bottom. Nothing runs earlier than it reads, the way a SQL WHERE runs before the SELECT written above it. And it means every chapter in this book can teach one verb at a time, because a verb never needs to know which verbs came before it.

There is one more fact about the sentence you just read, and it can wait for its own chapter: nothing ran until the page asked to see the answer. A pipeline is a plan, not a command, and the chapter on laziness shows what that buys you.

1.5 The first refusal, met on purpose

You will mistype something within the hour, so see now what that costs. A keep needs a question, and a column on its own is not one:

collect(sales |> keep(revenue))
Error:
! 
illegal: `keep where` needs a question that is either yes or no, and this is a number. Compare it to something: `is`, `>`, `<`, or `in {...}`
  |
2 |   then keep where [revenue]
  |                    ^^^^^^^
try:
    collect(sales >> keep(col.revenue))
except GodError as refusal:
    print(refusal)

illegal: `keep where` needs a question that is either yes or no, and this is a number. Compare it to something: `is`, `>`, `<`, or `in {...}`
  |
2 |   then keep where [revenue]
  |                    ^^^^^^^

Nothing ran and nothing broke. The message says what stopped and what to write instead, and every refusal in this book is held to that standard.

1.6 How to read the rest

Read the tab for the language you write. Glance at the others when you are curious. The explanation under a tabset is written once and covers every tab.

The next page lays out the whole vocabulary at once, so you can see how little there is to meet. The teaching chapters after it take the verbs one or two at a time. They share one shape, declared on the part page you just passed: a question, the verb that answers it, what the verb can vary, then what travels with it and what it refuses. They are short. By the end of them you can answer real questions about a table, and the part’s practice chapters will ask you to prove it.