10  Reading practice

Can you read a pipeline you have never seen? That was the preface’s day-one promise, and this chapter is where it is tested rather than repeated. Eight questions follow, each answered by one sentence built only from the words of this part. Read each pipeline aloud before you look at its table, then check the table against what you said. Nothing here is new; that is the point.

10.1 Which orders were larger than usual?

run('
  sales
    then keep where ([revenue] >= 200)
')
date region product quantity revenue cost
2026-01-09 East Widget 8 200 80
2026-01-26 West Gadget 5 300 200
2026-03-03 East Doohickey 5 200 125
2026-05-06 North Gadget 4 240 160
sales |> keep(revenue >= 200)
date region product quantity revenue cost
2026-01-09 East Widget 8 200 80
2026-01-26 West Gadget 5 300 200
2026-03-03 East Doohickey 5 200 125
2026-05-06 North Gadget 4 240 160
sales >> keep(col.revenue >= 200)
date region product quantity revenue cost
2026-01-09 East Widget 8 200 80
2026-01-26 West Gadget 5 300 200
2026-03-03 East Doohickey 5 200 125
2026-05-06 North Gadget 4 240 160

sales then keep where [revenue] >= 200

Four orders survive, and every region appears at least once, so size is not a story about one place.

10.2 What did each region sell, product by product?

run('
  sales
    then summarize [orders] as row_count(), [revenue] as total([revenue])
         by [region, product]
')
region product orders revenue
East Doohickey 1 200
East Gadget 2 180
East Sprocket 1 150
East Widget 1 200
North Doohickey 1 80
North Gadget 1 240
North Widget 2 225
West Doohickey 1 120
West Gadget 1 300
West Sprocket 1 90
West Widget 3 275
sales |>
  summarize(orders = row_count(), revenue = total(revenue),
            by = c(region, product))
region product orders revenue
East Doohickey 1 200
East Gadget 2 180
East Sprocket 1 150
East Widget 1 200
North Doohickey 1 80
North Gadget 1 240
North Widget 2 225
West Doohickey 1 120
West Gadget 1 300
West Sprocket 1 90
West Widget 3 275
(sales
  >> summarize(orders = row_count(), revenue = total(col.revenue),
               by = [col.region, col.product]))
region product orders revenue
East Doohickey 1 200
East Gadget 2 180
East Sprocket 1 150
East Widget 1 200
North Doohickey 1 80
North Gadget 1 240
North Widget 2 225
West Doohickey 1 120
West Gadget 1 300
West Sprocket 1 90
West Widget 3 275

sales then summarize [orders] as row_count(), [revenue] as total([revenue]) by [region, product]

Every combination that occurred gets a row, and one combination is absent: no North row mentions Sprocket, because the North never sold one. A grouped summary can only report what happened, which matters when a combination you expected is not there.

10.3 What does each product sell for?

run('
  sales
    then add [price] as ([revenue] / [quantity])
    then summarize [price] as average([price]), [sold] as total([quantity])
         by [product]
    then sort [price] descending
')
product price sold
Gadget 60 12
Doohickey 40 10
Widget 25 28
Sprocket 15 16
sales |>
  add(price = revenue / quantity) |>
  summarize(price = average(price), sold = total(quantity), by = product) |>
  sort(descending(price))
product price sold
Gadget 60 12
Doohickey 40 10
Widget 25 28
Sprocket 15 16
(sales
  >> add(price = col.revenue / col.quantity)
  >> summarize(price = average(col.price), sold = total(col.quantity),
               by = col.product)
  >> sort(descending(col.price)))
product price sold
Gadget 60 12
Doohickey 40 10
Widget 25 28
Sprocket 15 16

sales then add [price] as [revenue] / [quantity] then summarize [price] as average([price]), [sold] as total([quantity]) by [product] then sort [price] descending

Three verbs you know, in an order you can hear: work out each order’s unit price, average it per product, read the dearest first. The averages come out whole because each product held one price all year, which you can verify from the table faster than you can doubt it.

10.4 Which two orders cost the most to fulfill?

run('
  sales
    then sort [cost] descending
    then take 2
')
date region product quantity revenue cost
2026-01-26 West Gadget 5 300 200
2026-05-06 North Gadget 4 240 160
sales |> sort(descending(cost)) |> take(2)
date region product quantity revenue cost
2026-01-26 West Gadget 5 300 200
2026-05-06 North Gadget 4 240 160
sales >> sort(descending(col.cost)) >> take(2)
date region product quantity revenue cost
2026-01-26 West Gadget 5 300 200
2026-05-06 North Gadget 4 240 160

sales then sort [cost] descending then take 2

Both are Gadget orders, the product with the highest unit cost.

10.5 Who joined the survey before April?

run('
  survey
    then keep where ([joined] < "2026-04-01")
    then pick [name, region, joined]
')
name region joined
ana West 2026-01-12
ben East 2026-02-03
cal North 2026-02-27
dee West 2026-03-15
survey |> keep(joined < "2026-04-01") |> pick(name, region, joined)
name region joined
ana West 2026-01-12
ben East 2026-02-03
cal North 2026-02-27
dee West 2026-03-15
(survey
  >> keep(col.joined < "2026-04-01")
  >> pick(col.name, col.region, col.joined))
name region joined
ana West 2026-01-12
ben East 2026-02-03
cal North 2026-02-27
dee West 2026-03-15

survey then keep where [joined] < "2026-04-01" then pick [name, region, joined]

The comparison is on text, and it works because these dates are written largest unit first; the chapter on dates gives you the real date words for when text alone is not enough.

10.6 How many countries does each continent hold?

run('
  gapminder
    then keep where ([year] is 2007)
    then summarize [countries] as row_count() by [continent]
')
continent countries
Africa 52
Americas 25
Asia 33
Europe 30
Oceania 2
gapminder |>
  keep(year == 2007) |>
  summarize(countries = row_count(), by = continent)
continent countries
Africa 52
Americas 25
Asia 33
Europe 30
Oceania 2
(gapminder
  >> keep(col.year == 2007)
  >> summarize(countries = row_count(), by = col.continent))
continent countries
Africa 52
Americas 25
Asia 33
Europe 30
Oceania 2

gapminder then keep where [year] is 2007 then summarize [countries] as row_count() by [continent]

Same sentence shape as the sales questions, 1,704 rows instead of fifteen, and nothing about the spelling had to know the difference.

10.7 What was the cheapest order in each region?

run('
  sales
    then sort [cost]
    then take 1 by [region]
    then pick [region, product, cost]
')
region product cost
West Widget 20
North Widget 30
East Gadget 40
sales |> sort(cost) |> take(1, by = region) |> pick(region, product, cost)
region product cost
West Widget 20
North Widget 30
East Gadget 40
(sales
  >> sort(col.cost)
  >> take(1, by = col.region)
  >> pick(col.region, col.product, col.cost))
region product cost
West Widget 20
North Widget 30
East Gadget 40

sales then sort [cost] then take 1 by [region] then pick [region, product, cost]

The sort has no descending this time, so first means smallest. One word’s absence flipped the question, which is what reading aloud is for.

10.8 How much of its region is each order?

run('
  sales
    then add [share] as ([revenue] / total([revenue])) by [region]
    then sort [region], [share] descending
    then pick [region, product, revenue, share]
')
region product revenue share
East Widget 200 0.2739726
East Doohickey 200 0.2739726
East Sprocket 150 0.2054795
East Gadget 120 0.1643836
East Gadget 60 0.0821918
North Gadget 240 0.4403670
North Widget 150 0.2752294
North Doohickey 80 0.1467890
North Widget 75 0.1376147
West Gadget 300 0.3821656
West Widget 125 0.1592357
West Doohickey 120 0.1528662
West Widget 100 0.1273885
West Sprocket 90 0.1146497
West Widget 50 0.0636943
sales |>
  add(share = revenue / total(revenue), by = region) |>
  sort(region, descending(share)) |>
  pick(region, product, revenue, share)
region product revenue share
East Widget 200 0.2739726
East Doohickey 200 0.2739726
East Sprocket 150 0.2054795
East Gadget 120 0.1643836
East Gadget 60 0.0821918
North Gadget 240 0.4403670
North Widget 150 0.2752294
North Doohickey 80 0.1467890
North Widget 75 0.1376147
West Gadget 300 0.3821656
West Widget 125 0.1592357
West Doohickey 120 0.1528662
West Widget 100 0.1273885
West Sprocket 90 0.1146497
West Widget 50 0.0636943
(sales
  >> add(share = col.revenue / total(col.revenue), by = col.region)
  >> sort(col.region, descending(col.share))
  >> pick(col.region, col.product, col.revenue, col.share))
region product revenue share
East Widget 200 0.2739726
East Doohickey 200 0.2739726
East Sprocket 150 0.2054795
East Gadget 120 0.1643836
East Gadget 60 0.08219178
North Gadget 240 0.440367
North Widget 150 0.2752294
North Doohickey 80 0.146789
North Widget 75 0.1376147
West Gadget 300 0.3821656
West Widget 125 0.1592357
West Doohickey 120 0.1528662
West Widget 100 0.1273885
West Sprocket 90 0.1146497
West Widget 50 0.06369427

sales then add [share] as [revenue] / total([revenue]) by [region] then sort [region], [share] descending then pick [region, product, revenue, share]

Each region’s shares sum to one, and the biggest single share in the table sits in the North, where one Gadget order carries nearly half its region. This is the hardest sentence in the chapter, and it is made of a verb from adding a column and a word from the words that repeat.

10.9 What you just did

Eight questions, no new words. Every sentence was the everyday core doing its everyday work, and if you read them aloud, you have now heard the grammar more times than you will need to keep it. Reading was recognition. The next chapter asks for the harder thing, and it is the last thing this part asks of you: writing.