run('
sales
then pick [product, revenue]
')| product | revenue |
|---|---|
| Widget | 100 |
| Gadget | 120 |
| Doohickey | 120 |
| Widget | 150 |
| Widget | 200 |
| Gadget | 300 |
| Doohickey | 80 |
| Doohickey | 200 |
| Widget | 50 |
| Gadget | 240 |
| Sprocket | 150 |
| Sprocket | 90 |
| Widget | 75 |
| Gadget | 60 |
| Widget | 125 |
The table has six columns and your answer needs two of them. Which product, and what did it bring in? pick names the columns you want, and the rest do not travel any further down the pipeline.
run('
sales
then pick [product, revenue]
')| product | revenue |
|---|---|
| Widget | 100 |
| Gadget | 120 |
| Doohickey | 120 |
| Widget | 150 |
| Widget | 200 |
| Gadget | 300 |
| Doohickey | 80 |
| Doohickey | 200 |
| Widget | 50 |
| Gadget | 240 |
| Sprocket | 150 |
| Sprocket | 90 |
| Widget | 75 |
| Gadget | 60 |
| Widget | 125 |
sales |> pick(product, revenue)| product | revenue |
|---|---|
| Widget | 100 |
| Gadget | 120 |
| Doohickey | 120 |
| Widget | 150 |
| Widget | 200 |
| Gadget | 300 |
| Doohickey | 80 |
| Doohickey | 200 |
| Widget | 50 |
| Gadget | 240 |
| Sprocket | 150 |
| Sprocket | 90 |
| Widget | 75 |
| Gadget | 60 |
| Widget | 125 |
sales >> pick(col.product, col.revenue)| product | revenue |
|---|---|
| Widget | 100 |
| Gadget | 120 |
| Doohickey | 120 |
| Widget | 150 |
| Widget | 200 |
| Gadget | 300 |
| Doohickey | 80 |
| Doohickey | 200 |
| Widget | 50 |
| Gadget | 240 |
| Sprocket | 150 |
| Sprocket | 90 |
| Widget | 75 |
| Gadget | 60 |
| Widget | 125 |
sales then pick [product, revenue]
All fifteen rows are still here. keep and pick are the two axes of the same act of narrowing. keep decides which rows survive and never touches the columns; pick decides which columns survive and never touches the rows. Between them you can cut any table down to exactly the piece your question is about, and the two read as what they are, one verb per axis.
When the columns you want outnumber the ones you do not, name the leavers. Wrap the names in all_but:
run('
sales
then pick all_but [cost, region]
')| date | product | quantity | revenue |
|---|---|---|---|
| 2025-11-03 | Widget | 4 | 100 |
| 2025-11-17 | Gadget | 2 | 120 |
| 2025-12-05 | Doohickey | 3 | 120 |
| 2025-12-22 | Widget | 6 | 150 |
| 2026-01-09 | Widget | 8 | 200 |
| 2026-01-26 | Gadget | 5 | 300 |
| 2026-02-14 | Doohickey | 2 | 80 |
| 2026-03-03 | Doohickey | 5 | 200 |
| 2026-04-11 | Widget | 2 | 50 |
| 2026-05-06 | Gadget | 4 | 240 |
| 2026-06-19 | Sprocket | 10 | 150 |
| 2026-07-08 | Sprocket | 6 | 90 |
| 2026-08-24 | Widget | 3 | 75 |
| 2026-09-15 | Gadget | 1 | 60 |
| 2026-10-02 | Widget | 5 | 125 |
sales |> pick(all_but(cost, region))| date | product | quantity | revenue |
|---|---|---|---|
| 2025-11-03 | Widget | 4 | 100 |
| 2025-11-17 | Gadget | 2 | 120 |
| 2025-12-05 | Doohickey | 3 | 120 |
| 2025-12-22 | Widget | 6 | 150 |
| 2026-01-09 | Widget | 8 | 200 |
| 2026-01-26 | Gadget | 5 | 300 |
| 2026-02-14 | Doohickey | 2 | 80 |
| 2026-03-03 | Doohickey | 5 | 200 |
| 2026-04-11 | Widget | 2 | 50 |
| 2026-05-06 | Gadget | 4 | 240 |
| 2026-06-19 | Sprocket | 10 | 150 |
| 2026-07-08 | Sprocket | 6 | 90 |
| 2026-08-24 | Widget | 3 | 75 |
| 2026-09-15 | Gadget | 1 | 60 |
| 2026-10-02 | Widget | 5 | 125 |
sales >> pick(all_but(col.cost, col.region))| date | product | quantity | revenue |
|---|---|---|---|
| 2025-11-03 | Widget | 4 | 100 |
| 2025-11-17 | Gadget | 2 | 120 |
| 2025-12-05 | Doohickey | 3 | 120 |
| 2025-12-22 | Widget | 6 | 150 |
| 2026-01-09 | Widget | 8 | 200 |
| 2026-01-26 | Gadget | 5 | 300 |
| 2026-02-14 | Doohickey | 2 | 80 |
| 2026-03-03 | Doohickey | 5 | 200 |
| 2026-04-11 | Widget | 2 | 50 |
| 2026-05-06 | Gadget | 4 | 240 |
| 2026-06-19 | Sprocket | 10 | 150 |
| 2026-07-08 | Sprocket | 6 | 90 |
| 2026-08-24 | Widget | 3 | 75 |
| 2026-09-15 | Gadget | 1 | 60 |
| 2026-10-02 | Widget | 5 | 125 |
It is the same verb, deliberately. Choosing columns is choosing columns whichever way you say which ones. There is no second verb called drop here, for the same reason there was no second verb for dropping rows in the last chapter. all_but is a marker on the list, not a new idea, and it is the same word in both languages and in the text form.
pick also settles column order, because the order you name the columns is the order the table comes back in. There is no separate rearranging verb: if you want revenue first, ask for it first. One small consequence follows. A pick is a statement about the table it receives, so a column a previous step made is as pickable as one the table arrived with. And a column you picked away is gone for every later step. Each step sees only the table the last step handed it.
Naming columns one at a time works until the table is wide, or until next month’s file arrives with two more columns than this month’s. There is a second form, pick where, that chooses columns by a question about them instead of by name. It reuses the condition language you met in the last chapter, and it has a part of its own near the end of the book. The interesting questions there deserve more than a footnote: what is the name shaped like, and what kind of value does it hold? Until then, every table in this part is narrow enough to name.
all_but inverts any list of columns, here and anywhere a list of columns appears; you will meet it again in lengthen.where works inside this verb too. A pick where joins and negates its conditions the way keep does, asking about a column’s name and kind rather than its values, in choosing columns without naming them.A misspelled column is caught before anything runs, and the refusal does not stop at no:
collect(sales |> pick(product, reveune))Error:
!
illegal: there is no column called `reveune`. Did you mean `revenue`? The table has: date, region, product, quantity, revenue, cost
|
2 | then pick [product, reveune]
| ^^^^^^^
try:
collect(sales >> pick(col.product, col.reveune))
except GodError as refusal:
print(refusal)
illegal: there is no column called `reveune`. Did you mean `revenue`? The table has: date, region, product, quantity, revenue, cost
|
2 | then pick [product, reveune]
| ^^^^^^^
The message names the nearest real column and then lists everything the table has. The second most common cause of this error is not a typo at all: it is being wrong about which table you are holding. Both repairs are on the screen, and neither requires opening another window.
A whole table where a column belongs is answered before a sentence is even built, and the two languages get there differently. Python is handed the frame itself, so it says what it was given; R is handed the name products unevaluated, so the grammar looks for a column of that name and reports what the table actually has.
collect(sales |> pick(products))Error:
!
illegal: there is no column called `products`. Did you mean `product`? The table has: date, region, product, quantity, revenue, cost
|
2 | then pick [products]
| ^^^^^^^^
try:
collect(sales >> pick(products))
except GodError as refusal:
print(refusal)`pick` names a column, and this is a whole table. A column of it is written `col.name`
Both are GodError, which is the one exception either language asks you to catch. That matters because the mistake above is caught by the binding rather than by the engine, and a refusal you cannot catch the same way as every other refusal would be a second thing to remember for one idea.