16  Saying what something is

You can see the digits in the column, so why refuse to count them? Conversions are always written out. Nothing here converts a column on your behalf, because a column that quietly changes what it holds is the hardest kind of mistake to find later.

run('
  messy
    then add [word] as to_text([n])
    then add [digits] as characters([word])
')
raw n word digits
ann marie 7 7 1
bob 99 99 2
messy |> add(word = to_text(n)) |> add(digits = characters(word))
raw n word digits
ann marie 7 7 1
bob 99 99 2
messy >> add(word = to_text(col.n)) >> add(digits = characters(col.word))
raw n word digits
ann marie 7 7 1
bob 99 99 2

messy then add [word] as to_text([n]) then add [digits] as characters([word])

Every conversion begins to_, and nothing else does: to_number, to_text, to_date. Ask for the characters in a number without converting it first and you are told which conversion you wanted:

collect(messy |> add(size = characters(n)))
Error:
! 
illegal: `characters` counts the characters in text, and this is a number. Convert it first with `to_text(...)`
  |
2 |   then add [size] as characters([n])
  |                                  ^
try:
    collect(messy >> add(size = characters(col.n)))
except GodError as refusal:
    print(refusal)

illegal: `characters` counts the characters in text, and this is a number. Convert it first with `to_text(...)`
  |
2 |   then add [size] as characters([n])
  |                                  ^

16.1 The three conversions

to_text, to_number and to_date are the whole set. Each one says what it makes rather than what it takes, so there is nothing to remember about direction.

There were four of these, and the fourth was a mistake of the kind that hides. to_whole turned a number into a whole one, which sounds like a conversion and is not: this grammar has one kind of number, so to_whole converted nothing. It was a rounding wearing a conversion’s name, and because the name said nothing about direction, no reader could tell whether -5.5 came back as -5 or as -6. It is two words now, round_below and round_above, and they are in the chapter on adding a column, where the arithmetic lives.

run('
  messy
    then add [as_text] as to_text([n]), [half] as to_number([n]) / 2
    then pick [n, as_text, half]
')
n as_text half
7 7 3.5
99 99 49.5
messy |>
  add(as_text = to_text(n), half = to_number(n) / 2) |>
  pick(n, as_text, half)
n as_text half
7 7 3.5
99 99 49.5
(messy
  >> add(as_text = to_text(col.n), half = to_number(col.n) / 2)
  >> pick(col.n, col.as_text, col.half))
n as_text half
7 7 3.5
99 99 49.5

A conversion that cannot be made is the one answer this grammar does not own. to_number("abc") goes to the engine underneath, and engines disagree: one stops with an error, another hands back a missing value. It is one of the few places where the answer depends on where the pipeline ran, and the honest sentence is exactly that one.

16.2 Between two ends

between is not a conversion, and it is here because it is the other word you reach for when a column has just become a number. It asks whether a value falls in a range.

run('
  sales
    then keep where between([revenue], 150, 400)
    then sort [revenue]
')
date region product quantity revenue cost
2025-12-22 North Widget 6 150 60
2026-06-19 East Sprocket 10 150 120
2026-01-09 East Widget 8 200 80
2026-03-03 East Doohickey 5 200 125
2026-05-06 North Gadget 4 240 160
2026-01-26 West Gadget 5 300 200
sales |> keep(between(revenue, 150, 400)) |> sort(revenue)
date region product quantity revenue cost
2025-12-22 North Widget 6 150 60
2026-06-19 East Sprocket 10 150 120
2026-01-09 East Widget 8 200 80
2026-03-03 East Doohickey 5 200 125
2026-05-06 North Gadget 4 240 160
2026-01-26 West Gadget 5 300 200
sales >> keep(between(col.revenue, 150, 400)) >> sort(col.revenue)
date region product quantity revenue cost
2025-12-22 North Widget 6 150 60
2026-06-19 East Sprocket 10 150 120
2026-01-09 East Widget 8 200 80
2026-03-03 East Doohickey 5 200 125
2026-05-06 North Gadget 4 240 160
2026-01-26 West Gadget 5 300 200

Both ends are included, the way SQL and dplyr both have it, so nothing here needs checking against a manual.

to_date is the conversion this chapter has not run. It reads a date out of text, and the next chapter opens with it.