Skip to contents

This vignette covers the dplyr verbs that tidymatrix supports and how they keep the matrix in sync with its row and column metadata. For a quick overview of the whole package, start with the Getting started vignette.

library(tidymatrix)
library(dplyr, warn.conflicts = FALSE)

tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

big5 is a simulated personality survey: 400 respondents (rows) answered 30 questionnaire items (columns) on a 1–5 scale. See ?big5 for details.

Activation

A tidymatrix has three parts, and exactly one of them is active. dplyr verbs act on the active part:

Active Verbs act on Matrix follows by
rows row metadata (big5_respondents) subsetting/reordering its rows
columns column metadata (big5_items) subsetting/reordering its columns
matrix the matrix itself — (used by matrix operations)
tm |> activate(rows) |> active()
#> [1] "rows"
tm |> activate(columns)
#> # A tidymatrix: 400 x 30 matrix
#> # Active: columns
#> #
#> # Row data: 400 rows x 8 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (columns):
#>   item_id        trait reversed                             item_text position
#> 1      E1 Extraversion    FALSE     I feel comfortable around people.        1
#> 2      E2 Extraversion    FALSE                I start conversations.        6
#> 3      E3 Extraversion    FALSE I enjoy being part of a lively crowd.       11
#> 4      E4 Extraversion    FALSE                I make friends easily.       16
#> 5      E5 Extraversion     TRUE             I keep in the background.       21
#> 6      E6 Extraversion     TRUE    I have little to say to strangers.       26

Activation is sticky: it stays until you activate something else, so a pipeline can do several things to the rows before switching to the columns.

Filtering

filter() keeps the rows (or columns) of the metadata that match a condition and drops the corresponding rows (or columns) of the matrix.

Respondents who completed the survey in under 3.5 minutes did not read the questions. Let’s remove them:

tm_clean <- tm |>
  activate(rows) |>
  filter(completion_min > 3.5)

nrow(tm$matrix)
#> [1] 400
nrow(tm_clean$matrix)
#> [1] 388

Any expression that works in dplyr::filter() works here, including several conditions and helper functions:

tm_clean |>
  activate(rows) |>
  filter(age >= 30, age < 40, education %in% c("Master", "PhD"))
#> # A tidymatrix: 22 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 22 rows x 8 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#>   respondent_id age gender education   occupation country life_satisfaction
#> 1          R002  32 Female    Master       Office      FI                 4
#> 2          R021  39   Male    Master Professional      LV                 8
#> 3          R024  34 Female    Master     Creative      FI                 3
#> 4          R046  31   Male       PhD Professional      EE                 5
#> 5          R047  32 Female    Master      Student      LV                 6
#> 6          R055  33   Male    Master Professional      FI                 3
#>   completion_min
#> 1            6.2
#> 2           10.3
#> 3            9.8
#> 4            6.2
#> 5           10.6
#> 6            5.2

Filtering columns works the same way. Keep only the openness items that are not reverse-keyed:

tm_clean |>
  activate(columns) |>
  filter(trait == "Openness", !reversed)
#> # A tidymatrix: 388 x 4 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 8 columns
#> # Column data: 4 rows x 5 columns
#> #
#> # Active data (columns):
#>   item_id    trait reversed                                  item_text position
#> 1      O1 Openness    FALSE                I have a vivid imagination.        5
#> 2      O2 Openness    FALSE     I enjoy thinking about abstract ideas.       10
#> 3      O3 Openness    FALSE I like visiting art galleries and museums.       15
#> 4      O4 Openness    FALSE       I am quick to understand new things.       20

Adding and changing metadata

mutate() adds or changes metadata columns. It never touches the matrix, so the dimensions stay the same.

tm_clean <- tm_clean |>
  activate(rows) |>
  mutate(
    age_group = cut(
      age,
      breaks = c(17, 29, 49, 64, Inf),
      labels = c("18-29", "30-49", "50-64", "65+")
    ),
    satisfied = life_satisfaction >= 7
  )

tm_clean
#> # A tidymatrix: 388 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#>   respondent_id age gender education     occupation country life_satisfaction
#> 1          R001  52 Female Secondary Service/Manual      EE                 8
#> 2          R002  32 Female    Master         Office      FI                 4
#> 3          R003  66   Male  Bachelor         Office      EE                 6
#> 4          R004  42   Male  Bachelor   Professional      EE                 9
#> 5          R005  48 Female Secondary         Office      FI                 8
#> 6          R006  62 Female     Basic Service/Manual      EE                 4
#>   completion_min age_group satisfied
#> 1           19.8     50-64      TRUE
#> 2            6.2     30-49     FALSE
#> 3            6.3       65+     FALSE
#> 4            9.7     30-49      TRUE
#> 5           14.2     30-49      TRUE
#> 6           11.6     50-64     FALSE
tm_clean <- tm_clean |>
  activate(columns) |>
  mutate(label = paste0(item_id, if_else(reversed, " (R)", "")))

tm_clean |>
  activate(columns) |>
  pull(label)
#>  [1] "E1"     "E2"     "E3"     "E4"     "E5 (R)" "E6 (R)" "A1"     "A2"    
#>  [9] "A3"     "A4"     "A5 (R)" "A6 (R)" "C1"     "C2"     "C3"     "C4"    
#> [17] "C5 (R)" "C6 (R)" "N1"     "N2"     "N3"     "N4"     "N5 (R)" "N6 (R)"
#> [25] "O1"     "O2"     "O3"     "O4"     "O5 (R)" "O6 (R)"

Selecting, renaming and relocating metadata columns

These verbs change which metadata columns there are and in which order. They do not change the matrix.

tm_clean |>
  activate(rows) |>
  select(respondent_id, age, gender, education)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 388 rows x 4 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#>   respondent_id age gender education
#> 1          R001  52 Female Secondary
#> 2          R002  32 Female    Master
#> 3          R003  66   Male  Bachelor
#> 4          R004  42   Male  Bachelor
#> 5          R005  48 Female Secondary
#> 6          R006  62 Female     Basic

tm_clean |>
  activate(columns) |>
  rename(text = item_text) |>
  relocate(label, .after = item_id)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (columns):
#>   item_id  label        trait reversed                                  text
#> 1      E1     E1 Extraversion    FALSE     I feel comfortable around people.
#> 2      E2     E2 Extraversion    FALSE                I start conversations.
#> 3      E3     E3 Extraversion    FALSE I enjoy being part of a lively crowd.
#> 4      E4     E4 Extraversion    FALSE                I make friends easily.
#> 5      E5 E5 (R) Extraversion     TRUE             I keep in the background.
#> 6      E6 E6 (R) Extraversion     TRUE    I have little to say to strangers.
#>   position
#> 1        1
#> 2        6
#> 3       11
#> 4       16
#> 5       21
#> 6       26

pull() extracts one metadata column as a vector:

tm_clean |>
  activate(rows) |>
  pull(occupation) |>
  table()
#> 
#>       Creative         Office   Professional        Retired Service/Manual 
#>             33             89             80             25             86 
#>        Student     Unemployed 
#>             48             27

Reordering

arrange() sorts the metadata and reorders the matrix to match. The matrix columns of big5 are ordered by trait; to see the items in the order they were asked:

tm_ordered <- tm_clean |>
  activate(columns) |>
  arrange(position)

colnames(tm_ordered$matrix)[1:10]
#>  [1] "E1" "A1" "C1" "N1" "O1" "E2" "A2" "C2" "N2" "O2"

Sorting respondents by age, oldest first:

tm_by_age <- tm_clean |>
  activate(rows) |>
  arrange(desc(age))

tm_by_age |> activate(rows) |> select(respondent_id, age, occupation)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 388 rows x 3 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#>   respondent_id age occupation
#> 1          R029  73    Retired
#> 2          R238  73    Retired
#> 3          R117  72    Retired
#> 4          R264  72    Retired
#> 5          R052  70     Office
#> 6          R334  70    Retired
rownames(tm_by_age$matrix)[1:5]
#> [1] "R029" "R238" "R117" "R264" "R052"

Slicing

The slice() family selects rows or columns by position.

# first five respondents
tm_clean |> activate(rows) |> slice(1:5)
#> # A tidymatrix: 5 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 5 rows x 10 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#>   respondent_id age gender education     occupation country life_satisfaction
#> 1          R001  52 Female Secondary Service/Manual      EE                 8
#> 2          R002  32 Female    Master         Office      FI                 4
#> 3          R003  66   Male  Bachelor         Office      EE                 6
#> 4          R004  42   Male  Bachelor   Professional      EE                 9
#> 5          R005  48 Female Secondary         Office      FI                 8
#>   completion_min age_group satisfied
#> 1           19.8     50-64      TRUE
#> 2            6.2     30-49     FALSE
#> 3            6.3       65+     FALSE
#> 4            9.7     30-49      TRUE
#> 5           14.2     30-49      TRUE

# first three items
tm_clean |> activate(columns) |> slice_head(n = 3)
#> # A tidymatrix: 388 x 3 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 3 rows x 6 columns
#> #
#> # Active data (columns):
#>   item_id        trait reversed                             item_text position
#> 1      E1 Extraversion    FALSE     I feel comfortable around people.        1
#> 2      E2 Extraversion    FALSE                I start conversations.        6
#> 3      E3 Extraversion    FALSE I enjoy being part of a lively crowd.       11
#>   label
#> 1    E1
#> 2    E2
#> 3    E3

# a random 10% sample of respondents
set.seed(1)
tm_clean |> activate(rows) |> slice_sample(prop = 0.1)
#> # A tidymatrix: 38 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 38 rows x 10 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#>   respondent_id age gender education   occupation country life_satisfaction
#> 1          R331  56 Female    Master Professional      LT                 7
#> 2          R169  19 Female Secondary      Student      LV                 5
#> 3          R131  21 Female  Bachelor      Student      LV                 7
#> 4          R305  67   Male  Bachelor      Retired      SE                 6
#> 5          R275  30   Male  Bachelor      Student      FI                 7
#> 6          R190  44   Male Secondary       Office      EE                 7
#>   completion_min age_group satisfied
#> 1           12.9     50-64      TRUE
#> 2            8.9     18-29     FALSE
#> 3            5.5     18-29      TRUE
#> 4           10.3       65+     FALSE
#> 5            4.1     30-49      TRUE
#> 6            6.8     30-49      TRUE

Grouping and summarising

Grouping works as in dplyr, with one addition: summarise() also aggregates the matrix. Rows (or columns) in the same group are combined with a matrix function, mean() by default.

Summarising columns: one score per trait

Grouping the items by trait and summarising gives, for every respondent, the average answer per trait. Because the reverse-keyed items point in the opposite direction, we flip them first (see the Matrix operations vignette):

tm_scored <- tm_clean |>
  activate(rows) |>
  transform_matrix(\(x, flip) ifelse(flip, 6L - x, x), flip = reversed)

trait_scores <- tm_scored |>
  activate(columns) |>
  group_by(trait) |>
  summarise(n_items = n())

trait_scores
#> # A tidymatrix: 388 x 5 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 5 rows x 2 columns
#> #
#> # Active data (columns):
#>               trait n_items
#> 1     Agreeableness       6
#> 2 Conscientiousness       6
#> 3      Extraversion       6
#> 4       Neuroticism       6
#> 5          Openness       6
dim(trait_scores$matrix)
#> [1] 388   5
round(head(trait_scores$matrix), 2)
#>      [,1] [,2] [,3] [,4] [,5]
#> R001 3.33 1.50 3.33 3.00 2.83
#> R002 2.83 2.67 2.50 4.50 4.17
#> R003 3.33 2.17 2.67 3.17 1.83
#> R004 2.83 2.50 2.67 4.17 3.83
#> R005 4.83 3.67 1.33 3.33 2.50
#> R006 4.00 2.17 2.50 5.00 1.50

The result is again a tidymatrix, with one row per respondent and one column per trait. The respondent metadata is untouched.

Summarising rows: group profiles

Grouping respondents gives the average answer of each group to each item.

by_age <- tm_scored |>
  activate(rows) |>
  group_by(age_group) |>
  summarise(n = n(), mean_age = mean(age))

by_age
#> # A tidymatrix: 4 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 4 rows x 3 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#>   age_group   n mean_age
#> 1     18-29  83 24.61446
#> 2     30-49 178 38.64607
#> 3     50-64 102 55.74510
#> 4       65+  25 68.28000

Combining both steps gives a small age group × trait table:

age_trait <- by_age |>
  activate(columns) |>
  group_by(trait) |>
  summarise()

m <- age_trait$matrix
dimnames(m) <- list(
  age_trait$row_data$age_group,
  age_trait$col_data$trait
)
round(m, 2)
#>       Agreeableness Conscientiousness Extraversion Neuroticism Openness
#> 18-29          3.18              2.58         3.19        3.49     3.18
#> 30-49          3.31              2.79         3.10        3.15     3.21
#> 50-64          3.52              3.12         2.97        2.93     3.05
#> 65+            3.43              3.41         2.90        2.91     3.17

Conscientiousness and agreeableness increase with age and neuroticism decreases — a well-known pattern in personality research that is built into the simulated data.

Other aggregation functions

Use .matrix_fn to aggregate with something other than the mean, and .matrix_args to pass extra arguments to it:

tm_scored |>
  activate(rows) |>
  group_by(gender) |>
  summarise(n = n(), .matrix_fn = median)
#> # A tidymatrix: 3 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 3 rows x 2 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#>       gender   n
#> 1     Female 210
#> 2       Male 172
#> 3 Non-binary   6

Counting

count() is a shortcut for grouping, counting and summarising; tally() counts within existing groups.

tm_clean |>
  activate(rows) |>
  count(education)
#> # A tidymatrix: 5 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 5 rows x 2 columns
#> # Column data: 30 rows x 6 columns
#> #
#> # Active data (rows):
#>   education   n
#> 1     Basic  32
#> 2 Secondary 145
#> 3  Bachelor 131
#> 4    Master  67
#> 5       PhD  13

tm_clean |>
  activate(columns) |>
  group_by(trait, reversed) |>
  tally()
#> # A tidymatrix: 388 x 10 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 10 rows x 3 columns
#> #
#> # Active data (columns):
#>               trait reversed n
#> 1     Agreeableness    FALSE 4
#> 2     Agreeableness     TRUE 2
#> 3 Conscientiousness    FALSE 4
#> 4 Conscientiousness     TRUE 2
#> 5      Extraversion    FALSE 4
#> 6      Extraversion     TRUE 2

Putting it together

Verbs can be chained freely across rows and columns. The following pipeline takes the clean data, keeps working-age respondents with at least a Bachelor’s degree, keeps the conscientiousness and neuroticism items in questionnaire order, and adds a short label to each item:

tm_clean |>
  activate(rows) |>
  filter(
    !occupation %in% c("Student", "Retired"),
    education %in% c("Bachelor", "Master", "PhD")
  ) |>
  select(respondent_id, age, gender, education, occupation) |>
  activate(columns) |>
  filter(trait %in% c("Conscientiousness", "Neuroticism")) |>
  arrange(position) |>
  mutate(short = substr(item_text, 1, 25))
#> # A tidymatrix: 173 x 12 matrix
#> # Active: columns
#> #
#> # Row data: 173 rows x 5 columns
#> # Column data: 12 rows x 7 columns
#> #
#> # Active data (columns):
#>   item_id             trait reversed                     item_text position
#> 1      C1 Conscientiousness    FALSE I get chores done right away.        3
#> 2      N1       Neuroticism    FALSE    I get stressed out easily.        4
#> 3      C2 Conscientiousness    FALSE   I pay attention to details.        8
#> 4      N2       Neuroticism    FALSE         I worry about things.        9
#> 5      C3 Conscientiousness    FALSE  I like to follow a schedule.       13
#> 6      N3       Neuroticism    FALSE        My mood changes often.       14
#>   label                     short
#> 1    C1 I get chores done right a
#> 2    N1 I get stressed out easily
#> 3    C2 I pay attention to detail
#> 4    N2     I worry about things.
#> 5    C3 I like to follow a schedu
#> 6    N3    My mood changes often.

A note on stored analyses

Analysis functions such as compute_prcomp() store their result object in the tidymatrix. Verbs that remove, reorder or aggregate rows or columns (filter(), slice(), arrange(), joins, summarise(), count()) make those objects out of date, so they are dropped with a warning. The metadata columns the analysis created are kept. See the PCA and clustering vignette for details.

tm_pca <- tm_scored |>
  activate(columns) |>
  compute_prcomp(n_components = 2)

list_analyses(tm_pca)
#> [1] "column_pca"

tm_pca_filtered <- tm_pca |>
  activate(columns) |>
  filter(trait != "Openness")
#> Warning: Removed 1 stored analysis object(s) due to filter: column_pca
#> Metadata columns are preserved.

list_analyses(tm_pca_filtered)
#> character(0)