This vignette covers the dplyr verbs that tidymatrix supports and how they keep the matrix in sync with its row and column metadata. For a quick overview of the whole package, start with the Getting started vignette.
library(tidymatrix)
library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)big5 is a simulated personality survey: 400 respondents
(rows) answered 30 questionnaire items (columns) on a 1–5 scale. See
?big5 for details.
Activation
A tidymatrix has three parts, and exactly one of them is active. dplyr verbs act on the active part:
| Active | Verbs act on | Matrix follows by |
|---|---|---|
rows |
row metadata (big5_respondents) |
subsetting/reordering its rows |
columns |
column metadata (big5_items) |
subsetting/reordering its columns |
matrix |
the matrix itself | — (used by matrix operations) |
tm |> activate(rows) |> active()
#> [1] "rows"
tm |> activate(columns)
#> # A tidymatrix: 400 x 30 matrix
#> # Active: columns
#> #
#> # Row data: 400 rows x 8 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (columns):
#> item_id trait reversed item_text position
#> 1 E1 Extraversion FALSE I feel comfortable around people. 1
#> 2 E2 Extraversion FALSE I start conversations. 6
#> 3 E3 Extraversion FALSE I enjoy being part of a lively crowd. 11
#> 4 E4 Extraversion FALSE I make friends easily. 16
#> 5 E5 Extraversion TRUE I keep in the background. 21
#> 6 E6 Extraversion TRUE I have little to say to strangers. 26Activation is sticky: it stays until you activate something else, so a pipeline can do several things to the rows before switching to the columns.
Filtering
filter() keeps the rows (or columns) of the metadata
that match a condition and drops the corresponding rows (or columns) of
the matrix.
Respondents who completed the survey in under 3.5 minutes did not read the questions. Let’s remove them:
tm_clean <- tm |>
activate(rows) |>
filter(completion_min > 3.5)
nrow(tm$matrix)
#> [1] 400
nrow(tm_clean$matrix)
#> [1] 388Any expression that works in dplyr::filter() works here,
including several conditions and helper functions:
tm_clean |>
activate(rows) |>
filter(age >= 30, age < 40, education %in% c("Master", "PhD"))
#> # A tidymatrix: 22 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 22 rows x 8 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#> respondent_id age gender education occupation country life_satisfaction
#> 1 R002 32 Female Master Office FI 4
#> 2 R021 39 Male Master Professional LV 8
#> 3 R024 34 Female Master Creative FI 3
#> 4 R046 31 Male PhD Professional EE 5
#> 5 R047 32 Female Master Student LV 6
#> 6 R055 33 Male Master Professional FI 3
#> completion_min
#> 1 6.2
#> 2 10.3
#> 3 9.8
#> 4 6.2
#> 5 10.6
#> 6 5.2Filtering columns works the same way. Keep only the openness items that are not reverse-keyed:
tm_clean |>
activate(columns) |>
filter(trait == "Openness", !reversed)
#> # A tidymatrix: 388 x 4 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 8 columns
#> # Column data: 4 rows x 5 columns
#> #
#> # Active data (columns):
#> item_id trait reversed item_text position
#> 1 O1 Openness FALSE I have a vivid imagination. 5
#> 2 O2 Openness FALSE I enjoy thinking about abstract ideas. 10
#> 3 O3 Openness FALSE I like visiting art galleries and museums. 15
#> 4 O4 Openness FALSE I am quick to understand new things. 20Adding and changing metadata
mutate() adds or changes metadata columns. It never
touches the matrix, so the dimensions stay the same.
tm_clean <- tm_clean |>
activate(rows) |>
mutate(
age_group = cut(
age,
breaks = c(17, 29, 49, 64, Inf),
labels = c("18-29", "30-49", "50-64", "65+")
),
satisfied = life_satisfaction >= 7
)
tm_clean
#> # A tidymatrix: 388 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#> respondent_id age gender education occupation country life_satisfaction
#> 1 R001 52 Female Secondary Service/Manual EE 8
#> 2 R002 32 Female Master Office FI 4
#> 3 R003 66 Male Bachelor Office EE 6
#> 4 R004 42 Male Bachelor Professional EE 9
#> 5 R005 48 Female Secondary Office FI 8
#> 6 R006 62 Female Basic Service/Manual EE 4
#> completion_min age_group satisfied
#> 1 19.8 50-64 TRUE
#> 2 6.2 30-49 FALSE
#> 3 6.3 65+ FALSE
#> 4 9.7 30-49 TRUE
#> 5 14.2 30-49 TRUE
#> 6 11.6 50-64 FALSE
tm_clean <- tm_clean |>
activate(columns) |>
mutate(label = paste0(item_id, if_else(reversed, " (R)", "")))
tm_clean |>
activate(columns) |>
pull(label)
#> [1] "E1" "E2" "E3" "E4" "E5 (R)" "E6 (R)" "A1" "A2"
#> [9] "A3" "A4" "A5 (R)" "A6 (R)" "C1" "C2" "C3" "C4"
#> [17] "C5 (R)" "C6 (R)" "N1" "N2" "N3" "N4" "N5 (R)" "N6 (R)"
#> [25] "O1" "O2" "O3" "O4" "O5 (R)" "O6 (R)"Selecting, renaming and relocating metadata columns
These verbs change which metadata columns there are and in which order. They do not change the matrix.
tm_clean |>
activate(rows) |>
select(respondent_id, age, gender, education)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 388 rows x 4 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#> respondent_id age gender education
#> 1 R001 52 Female Secondary
#> 2 R002 32 Female Master
#> 3 R003 66 Male Bachelor
#> 4 R004 42 Male Bachelor
#> 5 R005 48 Female Secondary
#> 6 R006 62 Female Basic
tm_clean |>
activate(columns) |>
rename(text = item_text) |>
relocate(label, .after = item_id)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (columns):
#> item_id label trait reversed text
#> 1 E1 E1 Extraversion FALSE I feel comfortable around people.
#> 2 E2 E2 Extraversion FALSE I start conversations.
#> 3 E3 E3 Extraversion FALSE I enjoy being part of a lively crowd.
#> 4 E4 E4 Extraversion FALSE I make friends easily.
#> 5 E5 E5 (R) Extraversion TRUE I keep in the background.
#> 6 E6 E6 (R) Extraversion TRUE I have little to say to strangers.
#> position
#> 1 1
#> 2 6
#> 3 11
#> 4 16
#> 5 21
#> 6 26pull() extracts one metadata column as a vector:
Reordering
arrange() sorts the metadata and reorders the matrix to
match. The matrix columns of big5 are ordered by trait; to
see the items in the order they were asked:
tm_ordered <- tm_clean |>
activate(columns) |>
arrange(position)
colnames(tm_ordered$matrix)[1:10]
#> [1] "E1" "A1" "C1" "N1" "O1" "E2" "A2" "C2" "N2" "O2"Sorting respondents by age, oldest first:
tm_by_age <- tm_clean |>
activate(rows) |>
arrange(desc(age))
tm_by_age |> activate(rows) |> select(respondent_id, age, occupation)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 388 rows x 3 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#> respondent_id age occupation
#> 1 R029 73 Retired
#> 2 R238 73 Retired
#> 3 R117 72 Retired
#> 4 R264 72 Retired
#> 5 R052 70 Office
#> 6 R334 70 Retired
rownames(tm_by_age$matrix)[1:5]
#> [1] "R029" "R238" "R117" "R264" "R052"Slicing
The slice() family selects rows or columns by
position.
# first five respondents
tm_clean |> activate(rows) |> slice(1:5)
#> # A tidymatrix: 5 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 5 rows x 10 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#> respondent_id age gender education occupation country life_satisfaction
#> 1 R001 52 Female Secondary Service/Manual EE 8
#> 2 R002 32 Female Master Office FI 4
#> 3 R003 66 Male Bachelor Office EE 6
#> 4 R004 42 Male Bachelor Professional EE 9
#> 5 R005 48 Female Secondary Office FI 8
#> completion_min age_group satisfied
#> 1 19.8 50-64 TRUE
#> 2 6.2 30-49 FALSE
#> 3 6.3 65+ FALSE
#> 4 9.7 30-49 TRUE
#> 5 14.2 30-49 TRUE
# first three items
tm_clean |> activate(columns) |> slice_head(n = 3)
#> # A tidymatrix: 388 x 3 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 3 rows x 6 columns
#> #
#> # Active data (columns):
#> item_id trait reversed item_text position
#> 1 E1 Extraversion FALSE I feel comfortable around people. 1
#> 2 E2 Extraversion FALSE I start conversations. 6
#> 3 E3 Extraversion FALSE I enjoy being part of a lively crowd. 11
#> label
#> 1 E1
#> 2 E2
#> 3 E3
# a random 10% sample of respondents
set.seed(1)
tm_clean |> activate(rows) |> slice_sample(prop = 0.1)
#> # A tidymatrix: 38 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 38 rows x 10 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#> respondent_id age gender education occupation country life_satisfaction
#> 1 R331 56 Female Master Professional LT 7
#> 2 R169 19 Female Secondary Student LV 5
#> 3 R131 21 Female Bachelor Student LV 7
#> 4 R305 67 Male Bachelor Retired SE 6
#> 5 R275 30 Male Bachelor Student FI 7
#> 6 R190 44 Male Secondary Office EE 7
#> completion_min age_group satisfied
#> 1 12.9 50-64 TRUE
#> 2 8.9 18-29 FALSE
#> 3 5.5 18-29 TRUE
#> 4 10.3 65+ FALSE
#> 5 4.1 30-49 TRUE
#> 6 6.8 30-49 TRUEGrouping and summarising
Grouping works as in dplyr, with one addition:
summarise() also aggregates the matrix. Rows (or columns)
in the same group are combined with a matrix function,
mean() by default.
Summarising columns: one score per trait
Grouping the items by trait and summarising gives, for every respondent, the average answer per trait. Because the reverse-keyed items point in the opposite direction, we flip them first (see the Matrix operations vignette):
tm_scored <- tm_clean |>
activate(rows) |>
transform_matrix(\(x, flip) ifelse(flip, 6L - x, x), flip = reversed)
trait_scores <- tm_scored |>
activate(columns) |>
group_by(trait) |>
summarise(n_items = n())
trait_scores
#> # A tidymatrix: 388 x 5 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 5 rows x 2 columns
#> #
#> # Active data (columns):
#> trait n_items
#> 1 Agreeableness 6
#> 2 Conscientiousness 6
#> 3 Extraversion 6
#> 4 Neuroticism 6
#> 5 Openness 6
dim(trait_scores$matrix)
#> [1] 388 5
round(head(trait_scores$matrix), 2)
#> [,1] [,2] [,3] [,4] [,5]
#> R001 3.33 1.50 3.33 3.00 2.83
#> R002 2.83 2.67 2.50 4.50 4.17
#> R003 3.33 2.17 2.67 3.17 1.83
#> R004 2.83 2.50 2.67 4.17 3.83
#> R005 4.83 3.67 1.33 3.33 2.50
#> R006 4.00 2.17 2.50 5.00 1.50The result is again a tidymatrix, with one row per respondent and one column per trait. The respondent metadata is untouched.
Summarising rows: group profiles
Grouping respondents gives the average answer of each group to each item.
by_age <- tm_scored |>
activate(rows) |>
group_by(age_group) |>
summarise(n = n(), mean_age = mean(age))
by_age
#> # A tidymatrix: 4 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 4 rows x 3 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#> age_group n mean_age
#> 1 18-29 83 24.61446
#> 2 30-49 178 38.64607
#> 3 50-64 102 55.74510
#> 4 65+ 25 68.28000Combining both steps gives a small age group × trait table:
age_trait <- by_age |>
activate(columns) |>
group_by(trait) |>
summarise()
m <- age_trait$matrix
dimnames(m) <- list(
age_trait$row_data$age_group,
age_trait$col_data$trait
)
round(m, 2)
#> Agreeableness Conscientiousness Extraversion Neuroticism Openness
#> 18-29 3.18 2.58 3.19 3.49 3.18
#> 30-49 3.31 2.79 3.10 3.15 3.21
#> 50-64 3.52 3.12 2.97 2.93 3.05
#> 65+ 3.43 3.41 2.90 2.91 3.17Conscientiousness and agreeableness increase with age and neuroticism decreases — a well-known pattern in personality research that is built into the simulated data.
Other aggregation functions
Use .matrix_fn to aggregate with something other than
the mean, and .matrix_args to pass extra arguments to
it:
tm_scored |>
activate(rows) |>
group_by(gender) |>
summarise(n = n(), .matrix_fn = median)
#> # A tidymatrix: 3 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 3 rows x 2 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#> gender n
#> 1 Female 210
#> 2 Male 172
#> 3 Non-binary 6Counting
count() is a shortcut for grouping, counting and
summarising; tally() counts within existing groups.
tm_clean |>
activate(rows) |>
count(education)
#> # A tidymatrix: 5 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 5 rows x 2 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#> education n
#> 1 Basic 32
#> 2 Secondary 145
#> 3 Bachelor 131
#> 4 Master 67
#> 5 PhD 13
tm_clean |>
activate(columns) |>
group_by(trait, reversed) |>
tally()
#> # A tidymatrix: 388 x 10 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 10 rows x 3 columns
#> #
#> # Active data (columns):
#> trait reversed n
#> 1 Agreeableness FALSE 4
#> 2 Agreeableness TRUE 2
#> 3 Conscientiousness FALSE 4
#> 4 Conscientiousness TRUE 2
#> 5 Extraversion FALSE 4
#> 6 Extraversion TRUE 2Putting it together
Verbs can be chained freely across rows and columns. The following pipeline takes the clean data, keeps working-age respondents with at least a Bachelor’s degree, keeps the conscientiousness and neuroticism items in questionnaire order, and adds a short label to each item:
tm_clean |>
activate(rows) |>
filter(
!occupation %in% c("Student", "Retired"),
education %in% c("Bachelor", "Master", "PhD")
) |>
select(respondent_id, age, gender, education, occupation) |>
activate(columns) |>
filter(trait %in% c("Conscientiousness", "Neuroticism")) |>
arrange(position) |>
mutate(short = substr(item_text, 1, 25))
#> # A tidymatrix: 173 x 12 matrix
#> # Active: columns
#> #
#> # Row data: 173 rows x 5 columns
#> # Column data: 12 rows x 7 columns
#> #
#> # Active data (columns):
#> item_id trait reversed item_text position
#> 1 C1 Conscientiousness FALSE I get chores done right away. 3
#> 2 N1 Neuroticism FALSE I get stressed out easily. 4
#> 3 C2 Conscientiousness FALSE I pay attention to details. 8
#> 4 N2 Neuroticism FALSE I worry about things. 9
#> 5 C3 Conscientiousness FALSE I like to follow a schedule. 13
#> 6 N3 Neuroticism FALSE My mood changes often. 14
#> label short
#> 1 C1 I get chores done right a
#> 2 N1 I get stressed out easily
#> 3 C2 I pay attention to detail
#> 4 N2 I worry about things.
#> 5 C3 I like to follow a schedu
#> 6 N3 My mood changes often.A note on stored analyses
Analysis functions such as compute_prcomp() store their
result object in the tidymatrix. Verbs that remove, reorder or aggregate
rows or columns (filter(), slice(),
arrange(), joins, summarise(),
count()) make those objects out of date, so they are
dropped with a warning. The metadata columns the analysis created are
kept. See the PCA and clustering
vignette for details.
tm_pca <- tm_scored |>
activate(columns) |>
compute_prcomp(n_components = 2)
list_analyses(tm_pca)
#> [1] "column_pca"
tm_pca_filtered <- tm_pca |>
activate(columns) |>
filter(trait != "Openness")
#> Warning: Removed 1 stored analysis object(s) due to filter: column_pca
#> Metadata columns are preserved.
list_analyses(tm_pca_filtered)
#> character(0)