Appendix D — The verb you know

You arrive holding another tool’s word for it. This appendix is the index for that direction: find the name you know, read the god words beside it, and follow the link to the chapter that owns the idea. The other appendices run the opposite way, from god’s sentence outward.

The god column shows words, not whole pipelines, because most of these names map to a word plus a place to put it rather than to a verb alone.

D.1 From dplyr

you know god says the chapter
filter() keep Keeping rows
select(), select(-x) pick, all_but Choosing columns
mutate() add Adding a column
arrange(), desc() sort, descending Sorting and taking
group_by() + summarise() summarize ... by Summarizing by group
slice_head(n), slice_max() take, sort then take Sorting and taking
slice_tail(n) take_last, after a sort Sorting and taking
distinct() pick then drop_duplicates Renaming and repeats
rename(new = old) rename, same direction Renaming and repeats
pivot_longer() lengthen Columns into rows
pivot_wider() widen Rows into columns
left_join(), inner_join() join, unmatched Another table
semi_join(), anti_join() keep(matching(...)), with not Has a partner
bind_rows() add_rows Adding rows
complete(), expand() add_combinations, with by The rows that are not there
case_when() when ... otherwise One way or another
replace_values(), recode_values() look_up, with the ending written One way or another
coalesce() first_present Missing values
across(starts_with()) where, name, value One value per column
row_number(), min_rank() row_number, rank Each row a place
cumsum(), lag(), lead() running_total, previous, following The total so far
sd() standard_deviation Summarizing by group
paste(collapse = ), stringr::str_flatten() join_rows, after a sort Tidying text
slider::slide_dbl(), zoo::rollmean() rolling(average(...), n) The total so far
tidyr::replace_na(), tidyr::drop_na() fill_missing, drop_missing Missing values
tidyr::fill() latest, after a sort Missing values
show_query() show_as, with seven targets What it wrote
collect() collect, the same word Nothing runs until you ask

If you keep dplyr attached, four of the names in that column are also dplyr’s: collect, pick, rename and summarize. Three of them answer the same way whichever package you attached last, and pick is the one that does not. The names other packages take back is the whole of it.

D.2 From pandas

you know god says the chapter
query(), boolean mask keep Keeping rows
[[...]], drop(columns=) pick, all_but Choosing columns
assign() add Adding a column
sort_values(ascending=False) sort, descending Sorting and taking
groupby().agg() summarize ... by Summarizing by group
head(n), nlargest() take, sort then take Sorting and taking
tail(n) take_last, after a sort Sorting and taking
rename(columns={"old": "new"}) rename, the other direction Renaming and repeats
drop_duplicates() drop_duplicates, no subset Renaming and repeats
melt() lengthen Columns into rows
pivot_table(), pivot() widen Rows into columns
merge() join Another table
isin() filter against a table matching Has a partner
concat() add_rows Adding rows
reindex(MultiIndex.from_product()) add_combinations The rows that are not there
np.select, nested where when ... otherwise One way or another
.replace({...}), .map({...}) look_up, with the ending written One way or another
fillna(), dropna() fill_missing, drop_missing Missing values
ffill() latest, after a sort Missing values
cumsum(), shift() running_total, previous, following The total so far
std() standard_deviation Summarizing by group
.agg(", ".join), .str.cat() join_rows, after a sort Tidying text
sort_values(na_position=) sort ... missing first, or the default last Missing values
rolling(7).mean() rolling(average(...), 7) The total so far

D.3 From SQL

you know god says the chapter
WHERE keep Keeping rows
SELECT a, b pick Choosing columns
computed column in SELECT add Adding a column
ORDER BY ... DESC, LIMIT sort ... descending, take Sorting and taking
GROUP BY with aggregates summarize ... by Summarizing by group
SELECT DISTINCT pick then drop_duplicates Renaming and repeats
JOIN ... ON join ... by Another table
EXISTS, IN (SELECT ...) matching Has a partner
UNION ALL add_rows Adding rows
CROSS JOIN of two DISTINCTs, joined back add_combinations The rows that are not there
CASE WHEN ... ELSE ... END when ... otherwise One way or another
COALESCE first_present Missing values
RANK() OVER, ROW_NUMBER() rank, row_number Each row a place
SUM() OVER (ORDER BY), LAG, LEAD running_total, previous, following The total so far
STDDEV standard_deviation Summarizing by group
STRING_AGG, GROUP_CONCAT, LISTAGG join_rows, after a sort Tidying text
NULLS FIRST, NULLS LAST sort ... missing first, or the default last Missing values
LAST_VALUE(x IGNORE NULLS) OVER latest, after a sort Missing values
AVG() OVER (ROWS BETWEEN 6 PRECEDING AND CURRENT ROW) rolling(average(...), 7) The total so far

One absence from all three tables is deliberate. There is no row for your tool’s way of applying an arbitrary function to a column, because god has no word for it. The scalar vocabulary is closed, and the honest translation of apply is the exit the last table in Appendix E points at.