gapminder <- read.csv("../data/gapminder.csv")37 With gog
god manipulates tables. It does not draw them, and it is not going to, because drawing is a different job with a different vocabulary. A grammar that tried to do both would stop being small enough to hold in your head, which is the one thing this project will not trade.
That other grammar exists. gog is this project’s sibling: a grammar of graphics, by the same author, built on the same discipline and holding the same promise about day one and day two. Most real questions use both, because a question usually starts as a table that is the wrong shape and ends as a picture.
This chapter is one question answered end to end. Nothing else is loaded but pandas, and pandas is not a third grammar: it is how Python holds a table.
37.1 The table
Gapminder’s country panel: 142 countries, twelve observations each from 1952 to 2007, 1,704 rows in all. It is a real table rather than a demonstration, which is the point of using it here. The numbers are Gapminder’s, CC BY 4.0, and the file ships beside the book, so rendering this page needs no network.
It is also the one table this book reads from a file on the page, because a file is how your own table will arrive. One line in either language, and the panel the preface’s cast declared is an ordinary data frame:
gapminder = pd.read_csv("../data/gapminder.csv")From here it pipes like every table before it.
run('
gapminder
then take 5
')| country | continent | year | life | population | gdp |
|---|---|---|---|---|---|
| Afghanistan | Asia | 1952 | 28.801 | 8425333 | 779.4453 |
| Afghanistan | Asia | 1957 | 30.332 | 9240934 | 820.8530 |
| Afghanistan | Asia | 1962 | 31.997 | 10267083 | 853.1007 |
| Afghanistan | Asia | 1967 | 34.020 | 11537966 | 836.1971 |
| Afghanistan | Asia | 1972 | 36.088 | 13079460 | 739.9811 |
gapminder |> take(5)| country | continent | year | life | population | gdp |
|---|---|---|---|---|---|
| Afghanistan | Asia | 1952 | 28.801 | 8425333 | 779.4453 |
| Afghanistan | Asia | 1957 | 30.332 | 9240934 | 820.8530 |
| Afghanistan | Asia | 1962 | 31.997 | 10267083 | 853.1007 |
| Afghanistan | Asia | 1967 | 34.020 | 11537966 | 836.1971 |
| Afghanistan | Asia | 1972 | 36.088 | 13079460 | 739.9811 |
gapminder >> take(5)| country | continent | year | life | population | gdp |
|---|---|---|---|---|---|
| Afghanistan | Asia | 1952 | 28.801 | 8425333 | 779.4453 |
| Afghanistan | Asia | 1957 | 30.332 | 9240934 | 820.853 |
| Afghanistan | Asia | 1962 | 31.997 | 10267083 | 853.1007 |
| Afghanistan | Asia | 1967 | 34.02 | 11537966 | 836.1971 |
| Afghanistan | Asia | 1972 | 36.088 | 13079460 | 739.9811 |
Five rows of 1,704. Every table on this page is shown that way, because a chapter that prints 1,704 rows is not showing you anything.
37.2 The question
Which countries gained the most life expectancy, and what did the world look like while it happened?
The first half is the table’s job. Each country has twelve rows and the answer needs two of them, the earliest and the latest, which is what first and last are for. They answer by position, so a sort in front of them is what decides which position that is.
gains <- gapminder |>
sort(year) |>
summarize(started = first(life), ended = last(life),
by = c(country, continent)) |>
add(gained = ended - started) |>
sort(descending(gained))
gains |> take(5)| country | continent | started | ended | gained |
|---|---|---|---|---|
| Oman | Asia | 37.578 | 75.640 | 38.062 |
| Vietnam | Asia | 40.412 | 74.249 | 33.837 |
| Indonesia | Asia | 37.468 | 70.650 | 33.182 |
| Saudi Arabia | Asia | 39.875 | 72.777 | 32.902 |
| Libya | Africa | 42.723 | 73.952 | 31.229 |
gains = (gapminder
>> sort(col.year)
>> summarize(started = first(col.life), ended = last(col.life),
by = [col.country, col.continent])
>> add(gained = col.ended - col.started)
>> sort(descending(col.gained)))
gains >> take(5)| country | continent | started | ended | gained |
|---|---|---|---|---|
| Oman | Asia | 37.578 | 75.64 | 38.062 |
| Vietnam | Asia | 40.412 | 74.249 | 33.837 |
| Indonesia | Asia | 37.468 | 70.65 | 33.182 |
| Saudi Arabia | Asia | 39.875 | 72.777 | 32.902 |
| Libya | Africa | 42.723 | 73.952 | 31.229 |
Oman gained more than thirty years of life expectancy in half a century. That is one sentence of god: sort, summarize by two columns, subtract, sort again.
The other end of the same table is a different question and the same pipeline with one word changed.
gains |> sort(gained) |> take(5)| country | continent | started | ended | gained |
|---|---|---|---|---|
| Zimbabwe | Africa | 48.451 | 43.487 | -4.964 |
| Swaziland | Africa | 41.407 | 39.613 | -1.794 |
| Zambia | Africa | 42.038 | 42.384 | 0.346 |
| Lesotho | Africa | 42.138 | 42.592 | 0.454 |
| Botswana | Africa | 47.622 | 50.728 | 3.106 |
gains >> sort(col.gained) >> take(5)| country | continent | started | ended | gained |
|---|---|---|---|---|
| Zimbabwe | Africa | 48.451 | 43.487 | -4.964 |
| Swaziland | Africa | 41.407 | 39.613 | -1.794 |
| Zambia | Africa | 42.038 | 42.384 | 0.346 |
| Lesotho | Africa | 42.138 | 42.592 | 0.454 |
| Botswana | Africa | 47.622 | 50.728 | 3.106 |
37.3 When you need collect
Nothing above called collect, and the tables still appeared. Everyone asks about that once.
A pipeline is a plan, and it does not run when you write it. Printing one runs it. That is why every example so far has shown a table: this book prints each pipeline, so each one is asked for its answer.
collect is how you ask for the answer when something other than printing needs it. Handing a table to another function is exactly that case, and gog is about to be handed one.
top <- gains |> take(12) |> collect()
class(top)[1] "data.frame"
top = collect(gains >> take(12))
type(top).__name__'DataFrame'
So the rule is short. If you want to look at it, print it. If you want to use it, collect it.
37.4 The picture
Now gog. The table goes in, a mark and some channels say what to draw, and the sentence reads the way a god pipeline does. The table going in is top, the first twelve rows of the ranking you have already read. On the Python side every gog name arrives through the module, as gog.point and gog.col. The section near the end of this chapter says why, and why R needs no prefix at all.
data(top) + bar + x(gained) + y(country) + color(continent)(gog.data(top) + gog.bar + gog.x(gog.col.gained) + gog.y(gog.col.country)
+ gog.color(gog.col.continent))37.5 What the gain looked like while it happened
The ranking says who gained. The shape of the gain is a second god sentence, one summarize with two columns after by:
trends <- gapminder |>
summarize(life = average(life), by = c(continent, year)) |>
sort(year) |>
collect()
trends |> take(5)| continent | year | life |
|---|---|---|
| Africa | 1952 | 39.13550 |
| Asia | 1952 | 46.31439 |
| Europe | 1952 | 64.40850 |
| Oceania | 1952 | 69.25500 |
| Americas | 1952 | 53.27984 |
trends = collect(gapminder
>> summarize(life = average(col.life), by = [col.continent, col.year])
>> sort(col.year))
trends >> take(5)| continent | year | life |
|---|---|---|
| Africa | 1952 | 39.1355 |
| Asia | 1952 | 46.31439 |
| Europe | 1952 | 64.4085 |
| Oceania | 1952 | 69.255 |
| Americas | 1952 | 53.27984 |
Sixty rows came back, one per continent per year, and the first five are the five continents in 1952. A line per continent is then one gog sentence:
data(trends) + line + x(year) + y(life) + color(continent)(gog.data(trends) + gog.line + gog.x(gog.col.year) + gog.y(gog.col.life)
+ gog.color(gog.col.continent))Every continent climbs across the half century, and the distance between the top line and the bottom one is the story the rest of this page moves through.
37.6 What the whole world was doing
The bar chart answers the question. This next one is why the pairing exists: the same panel, unsummarized, with one more channel.
play turns a column into time. Every country is a point, GDP per person runs along the bottom on a log scale, and life expectancy up the side. The year becomes the animation rather than another axis.
data(gapminder) + point + x(gdp, scale = "log") + y(life) +
color(continent) + size(population) + play(year)(gog.data(gapminder) + gog.point + gog.x(gog.col.gdp, scale = "log")
+ gog.y(gog.col.life) + gog.color(gog.col.continent)
+ gog.size(gog.col.population) + gog.play(gog.col.year))That is the Rosling chart, and neither package needed a special case to draw it. play is a channel like any other, so the year binds the way a color does.
37.7 One panel per continent
facet splits a plot into panels, one per value of a column. It joins with | rather than +, because it acts on the whole plot rather than adding a channel to it. The moving chart above, five times:
data(gapminder) + point + x(gdp, scale = "log") + y(life) +
color(continent) + play(year) | facet(continent)(gog.data(gapminder) + gog.point + gog.x(gog.col.gdp, scale = "log")
+ gog.y(gog.col.life) + gog.color(gog.col.continent)
+ gog.play(gog.col.year)) | gog.facet(gog.col.continent)Five panels, one clock. The year advances everywhere at once, so a continent that stalls is read against the ones that do not.
37.8 A third position
z is the third position, and writing it is what turns a plot into a cube. The slice of one year is a god step:
latest <- gapminder |> keep(year == 2007) |> collect()
latest |> take(5)| country | continent | year | life | population | gdp |
|---|---|---|---|---|---|
| Afghanistan | Asia | 2007 | 43.828 | 31889923 | 974.5803 |
| Albania | Europe | 2007 | 76.423 | 3600523 | 5937.0295 |
| Algeria | Africa | 2007 | 72.301 | 33333216 | 6223.3675 |
| Angola | Africa | 2007 | 42.731 | 12420476 | 4797.2313 |
| Argentina | Americas | 2007 | 75.320 | 40301927 | 12779.3796 |
latest = collect(gapminder >> keep(col.year == 2007))
latest >> take(5)| country | continent | year | life | population | gdp |
|---|---|---|---|---|---|
| Afghanistan | Asia | 2007 | 43.828 | 31889923 | 974.5803 |
| Albania | Europe | 2007 | 76.423 | 3600523 | 5937.03 |
| Algeria | Africa | 2007 | 72.301 | 33333216 | 6223.367 |
| Angola | Africa | 2007 | 42.731 | 12420476 | 4797.231 |
| Argentina | Americas | 2007 | 75.32 | 40301927 | 12779.38 |
One row per country, 142 in all, and the cube binds three of the columns. space names the viewing angle, turn and tilt, for when the default angle hides what you came to see:
data(latest) + point + x(gdp, scale = "log") + y(life) + z(population) +
color(continent) + space(turn = -50, tilt = 15)(gog.data(latest) + gog.point + gog.x(gog.col.gdp, scale = "log")
+ gog.y(gog.col.life) + gog.z(gog.col.population)
+ gog.color(gog.col.continent) + gog.space(turn = -50, tilt = 15))And play is still a channel, so the cube moves the way the flat chart did, with nothing taught about the combination:
data(gapminder) + point + x(gdp, scale = "log") + y(life) + z(population) +
color(continent) + play(year)(gog.data(gapminder) + gog.point + gog.x(gog.col.gdp, scale = "log")
+ gog.y(gog.col.life) + gog.z(gog.col.population)
+ gog.color(gog.col.continent) + gog.play(gog.col.year))37.9 Handing the reader the choice
brush lights a region and pushes the rest back, and it removes nothing. Bounded in the sentence, it prints with the selection already made; written bare, as + brush, it waits for a reader to drag:
data(latest) + point + x(gdp, scale = "log") + y(life) +
color(continent) + brush(gdp, at = c(20000, 60000))(gog.data(latest) + gog.point + gog.x(gog.col.gdp, scale = "log")
+ gog.y(gog.col.life) + gog.color(gog.col.continent)
+ gog.brush(gog.col.gdp, at = [20000, 60000]))The countries above twenty thousand dollars sit almost entirely above seventy-five years, and everything else stays in view behind them. That is the working difference between the two grammars’ questions. god’s keep removes rows before the answer; gog’s brush highlights them inside it. The comparison above needs the context a keep would have thrown away.
37.10 What it replaced
Written in the tools this book keeps comparing itself against, the page above needs a table package for the summarizing and a plotting system on top of it. Each brings its own vocabulary, its own idea of what a column reference is, and its own rules about which arguments are quoted. The animation usually needs a third package, and the cube and the brush a fourth.
The saving is not the line count. It is that summarize and play come from the same design, so the second half of the analysis does not ask you to change how you are thinking. by means in gog what it means here. So does color.
37.11 The one thing to know about using both
In R, load them and go. R looks a column name up in your data before it looks anywhere else, and the two packages take care not to collide.
Python is a different shape. The rule there is short: bring god in whole, and reach gog through its name, which is what the examples above do.
from god import *
import gogThen col.life is god’s and gog.col.life is gog’s, and nothing is ambiguous.
Two things make that the right way round, rather than a preference for whichever book you happen to be reading. Two names are claimed by both: col, which is how each of them names a column, and median, a word both vocabularies wanted. Bringing in both whole would leave you holding gog’s pair and no warning.
The second reason would still apply if god did not exist. A grammar of graphics wants words the language has already spent. sum, max and several more are transforms in gog and builtins in Python, so taking gog whole takes those from you as well. god’s vocabulary overlaps none of Python’s, and that is what makes it the one that is safe to bring in whole.
So every gog name on this page carries the prefix, not only the two that collide. One rule is easier to keep than a list of which names happen to be free this year.
It reverses cleanly when a file is mostly pictures. Take gog whole, reach god through its name, and take back the builtins you use with from builtins import sum. What does not reverse is taking both whole, which is the one arrangement here that fails without saying so.
Appendix A says the same thing about every other name god takes over.
37.12 The question that needs two steps
Halfway between the table and the plot, a tempting shortcut appears: keep the countries above the average, in one step. keep decides one row at a time, so it cannot ask a question about a whole group, and the refusal says which two steps to write instead:
collect(gapminder |> keep(average(life) > 50))Error:
!
illegal: `keep` decides one row at a time, so it cannot ask a question about a whole group. Summarize first, then keep: `then summarize [n] as row_count() by [g] then keep where [n] > 5`
|
2 | then keep where (average([life]) > 50)
| ^^^^^^^^^^^^^^^^^^^^
try:
collect(gapminder >> keep(average(col.life) > 50))
except GodError as refusal:
print(refusal)
illegal: `keep` decides one row at a time, so it cannot ask a question about a whole group. Summarize first, then keep: `then summarize [n] as row_count() by [g] then keep where [n] > 5`
|
2 | then keep where (average([life]) > 50)
| ^^^^^^^^^^^^^^^^^^^^
The two-step spelling it offers is the shape this chapter’s own pipeline already has: the group answer first, and everything that reads it afterwards.
37.13 Where to go next
The graphics grammar has its own manual, at the gog book, written the way this one is. Every plot on the page was drawn by the engine while the page was built.
Read it in the order you would use them. You start in god, with a table that is the wrong shape. You end in gog, looking at the answer.
The table on this page is the Gapminder panel, in the excerpt that Jennifer Bryan’s gapminder R package made standard (Bryan, 2025). The figures are the Gapminder Foundation’s, CC BY 4.0, and the credit belongs in both places.