Data Receipt is for asking people for data in a shape you choose. You define a request (a set of columns, each with a type and rules), share its link, and every spreadsheet uploaded against it is validated before it is accepted. This package is the other end of that: it brings the accepted data into R.
Your key
Every call is authenticated with an API token that belongs to your
account. Create one on the site under Settings > API
tokens, then save it once with install = TRUE:
library(datareceipt)
datareceipt_api_key("12|dr_...", install = TRUE)This appends DATARECEIPT_API_KEY to your
.Renviron, so R finds it in every future session. Restart R
(or run readRenviron("~/.Renviron")) and check:
What have I asked for?
list_requests() is a tibble of your requests, newest
first, with how many submissions and rows each has received:
requests <- list_requests()
requestsEach request’s spec is in the columns list column, or
ask for it directly. This is the shape your data will come back in, so
it is worth a look before you start:
get_columns(requests$id[[1]])type is one of text, numeric,
integer, date, or boolean;
required means no empty cells were allowed;
min, max, and allowed_values are
the rules senders had to meet. on_break says what happens
to a value that breaks them: block stops the submission,
flag accepts the value and marks it (more on that
below).
The data
get_data() returns every accepted row across every
submission to a request, as one tibble. Pass the id, or paste the
request’s URL from your browser:
data <- get_data(requests$id[[1]])
dataColumns are typed from the spec: text is character, numbers are
double, whole numbers are integer, dates are Date, yes/no
is logical, and an empty cell is NA. The first columns say
where each row came from: submission_id, row
(its position within that submission), submitted_at,
sender_name, and sender_email. Leave the
sender columns out with include_sender = FALSE.
From here it is ordinary tidyverse work:
Flagged cells
A column set to flag accepts a value that breaks its
rules and keeps it as the sender typed it. In get_data()
those cells are cast where the text allows it, so an out-of-range
"200" in a whole-number column is still 200L,
and are NA where it does not, like "abc" in
the same column. get_data() tells you how many cells were
flagged, and get_flags() lists them: where each one is, the
text as sent, and the rule it broke.
get_flags(requests$id[[1]])One submission at a time
list_submissions() shows who sent what, how
(source: an uploaded xlsx or csv;
submissions from August 2026 may read paste,
typed, or form, from sender methods the site
no longer offers), and when, along with how many cells were flagged and
whether the request’s owner has since edited the submission on the site
(revised_at). get_submission() fetches one
submission’s rows in the same shape as get_data():
submissions <- list_submissions(requests$id[[1]])
submissions
get_submission(submissions$id[[1]])