Skip to contents

dplyr verbs change which rows and columns a tidymatrix has. Matrix operations change the values. As elsewhere in the package, the active component decides the direction: with rows active an operation works row by row, with columns active column by column, and with the matrix active on all values at once.

library(tidymatrix)
library(dplyr, warn.conflicts = FALSE)

tm <- tidymatrix(big5_responses, big5_respondents, big5_items)

We use the big5 personality survey: 400 respondents answering 30 items on a 1–5 scale (see ?big5).

Summary statistics into metadata

add_stats() computes statistics along the active dimension and stores them as metadata columns. Per-respondent statistics are a standard data quality check in surveys: someone who gives the same answer to every question has a standard deviation close to zero.

tm <- tm |>
  activate(rows) |>
  add_stats(mean, sd, .names = c("resp_mean", "resp_sd"))

tm |>
  activate(rows) |>
  arrange(resp_sd) |>
  select(respondent_id, resp_mean, resp_sd, completion_min)
#> # A tidymatrix: 400 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 400 rows x 4 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#>   respondent_id resp_mean   resp_sd completion_min
#> 1          R074  3.066667 0.3651484            2.0
#> 2          R307  3.900000 0.4025779            3.2
#> 3          R392  4.900000 0.4025779            1.5
#> 4          R170  3.966667 0.4138410            2.5
#> 5          R219  3.900000 0.5477226            2.2
#> 6          R352  4.866667 0.5713465            2.6

The respondents with the least variable answers also finished the survey in a couple of minutes. Sorting by completion time shows the other kind of careless respondent — those who answered at random, and so have a normal-looking standard deviation:

tm |>
  activate(rows) |>
  arrange(completion_min) |>
  select(respondent_id, resp_sd, completion_min)
#> # A tidymatrix: 400 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 400 rows x 3 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#>   respondent_id   resp_sd completion_min
#> 1          R392 0.4025779            1.5
#> 2          R282 1.3746473            1.7
#> 3          R393 1.2793677            1.7
#> 4          R074 0.3651484            2.0
#> 5          R368 1.4762487            2.0
#> 6          R219 0.5477226            2.2

We remove both groups and move on with the clean data:

tm <- tm |>
  activate(rows) |>
  filter(completion_min > 3.5)

Per-item statistics work the same way with columns active. Custom functions can be passed as a named list with .fns:

tm |>
  activate(columns) |>
  add_stats(.fns = list(
    item_mean = mean,
    item_sd = sd,
    pct_agree = \(x) mean(x >= 4)
  )) |>
  select(item_id, item_text, item_mean, item_sd, pct_agree)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (columns):
#>   item_id                             item_text item_mean  item_sd pct_agree
#> 1      E1     I feel comfortable around people.  2.951031 1.145793 0.3350515
#> 2      E2                I start conversations.  3.087629 1.117191 0.3762887
#> 3      E3 I enjoy being part of a lively crowd.  3.373711 1.123868 0.4664948
#> 4      E4                I make friends easily.  3.257732 1.163869 0.4355670
#> 5      E5             I keep in the background.  2.935567 1.191474 0.3221649
#> 6      E6    I have little to say to strangers.  3.298969 1.096529 0.4432990

Custom transformations with transform_matrix()

transform_matrix() applies any function to the matrix. What the function receives depends on the active component:

Active fn receives Example
matrix the whole matrix log, \(m) m / 4
rows one row at a time per-respondent ranks
columns one column at a time per-item standardisation

The function must return values of the same size, and the row and column names are preserved. Extra arguments are passed on to the function; with rows active they are evaluated in the column metadata, with columns active in the row metadata (use .env$x for a variable x that clashes with a metadata column).

Reverse-scoring items

A questionnaire usually measures each trait with several statements, and the answers are later combined, for example averaged, into one score per trait. Some of the statements are deliberately worded in the opposite direction. These are called reverse-keyed items. They keep people who tend to agree with everything (or tick the same box throughout) from automatically ending up with high scores.

In big5, two of the six items per trait are reverse-keyed. For example:

Trait Normally keyed Reverse-keyed
Extraversion “I start conversations.” “I keep in the background.”
Conscientiousness “I finish what I start.” “I put off important tasks.”
Neuroticism “I worry about things.” “I stay calm under pressure.”

Answering 5 (“strongly agree”) to “I keep in the background” indicates low extraversion. Averaged as they are, such answers would cancel out the normally keyed ones. So before the items are combined, the answers to reverse-keyed items are flipped: 5 becomes 1, 4 becomes 2, 3 stays 3. On a 1–5 scale that is simply 6 - x. Afterwards a high value always means a high trait level. This step is known as reverse-scoring.

Which columns to flip is stored in the column metadata. With rows active the function gets one respondent’s answers, one value per item. Extra arguments to transform_matrix() are passed on to the function, and they are evaluated with the column metadata as a data mask — just like the arguments of mutate(). So flip = reversed hands the function the reversed column, which lines up element by element with the answers:

tm_scored <- tm |>
  activate(rows) |>
  transform_matrix(\(x, flip) ifelse(flip, 6L - x, x), flip = reversed)

# E4 is positively keyed, E5 is reverse-keyed
tm$matrix[1:5, c("E4", "E5")]
#>      E4 E5
#> R001  3  3
#> R002  2  3
#> R003  4  4
#> R004  3  3
#> R005  1  5
tm_scored$matrix[1:5, c("E4", "E5")]
#>      E4 E5
#> R001  3  3
#> R002  2  3
#> R003  4  2
#> R004  3  3
#> R005  1  1

After scoring, items of the same trait correlate positively:

round(cor(tm$matrix[, paste0("E", 1:6)])[5:6, 1:4], 2)
#>       E1    E2    E3    E4
#> E5 -0.45 -0.56 -0.42 -0.50
#> E6 -0.41 -0.41 -0.34 -0.37
round(cor(tm_scored$matrix[, paste0("E", 1:6)])[5:6, 1:4], 2)
#>      E1   E2   E3   E4
#> E5 0.45 0.56 0.42 0.50
#> E6 0.41 0.41 0.34 0.37

Whole-matrix transformations

With the matrix active, the function gets the full matrix. Rescaling the 1–5 answers to a 0–100 scale:

tm_scored |>
  activate(matrix) |>
  transform_matrix(\(m) (m - 1) / 4 * 100) |>
  pull_active() |>
  head(3)
#>      E1 E2  E3 E4 E5 E6 A1  A2 A3  A4 A5 A6 C1 C2 C3 C4 C5 C6  N1 N2  N3  N4 N5
#> R001 75 50 100 50 50 25 50 100 50  75 25 50  0 25  0 25 25  0  50 50  50  25 50
#> R002 25 75  25 25 50 25 50  25 50  50 75 25 50 75 50 25 25 25 100 75 100 100 75
#> R003 25 25  50 75 25 50 50  50 75 100 25 50 25 50  0 50 25 25  75 75  50  25 25
#>      N6  O1 O2 O3  O4 O5  O6
#> R001 75  25 75  0  50 75  50
#> R002 75 100 50 75 100 50 100
#> R003 75   0 75  0  25  0  25

Column- and row-wise transformations

With columns active the function is applied to each item separately. For example, percentile ranks within each item:

tm_scored |>
  activate(columns) |>
  transform_matrix(\(x) rank(x) / length(x)) |>
  activate(matrix) |>
  pull_active() |>
  head(3) |>
  round(2)
#>        E1   E2   E3   E4   E5  E6   A1   A2   A3   A4   A5   A6   C1   C2   C3
#> R001 0.79 0.47 0.91 0.41 0.48 0.3 0.43 0.87 0.41 0.57 0.30 0.30 0.03 0.31 0.10
#> R002 0.24 0.76 0.14 0.17 0.48 0.3 0.43 0.10 0.41 0.27 0.86 0.11 0.37 0.84 0.61
#> R003 0.24 0.20 0.38 0.70 0.22 0.6 0.43 0.31 0.70 0.87 0.30 0.30 0.15 0.60 0.10
#>        C4  C5   C6   N1   N2   N3   N4   N5   N6   O1   O2   O3   O4   O5   O6
#> R001 0.19 0.4 0.08 0.45 0.52 0.48 0.20 0.52 0.57 0.19 0.60 0.08 0.34 0.77 0.54
#> R002 0.19 0.4 0.29 0.93 0.80 0.93 0.93 0.79 0.57 0.92 0.32 0.81 0.90 0.50 0.95
#> R003 0.46 0.4 0.29 0.72 0.80 0.48 0.20 0.24 0.57 0.05 0.60 0.08 0.13 0.05 0.26

Extra arguments are passed on to the function:

tm_scored |>
  activate(columns) |>
  transform_matrix(\(x) x / 4) |>
  activate(matrix) |>
  transform_matrix(round, digits = 1) |>
  pull_active() |>
  head(3)
#>       E1  E2  E3  E4  E5  E6  A1  A2  A3  A4  A5  A6  C1  C2  C3  C4  C5  C6
#> R001 1.0 0.8 1.2 0.8 0.8 0.5 0.8 1.2 0.8 1.0 0.5 0.8 0.2 0.5 0.2 0.5 0.5 0.2
#> R002 0.5 1.0 0.5 0.5 0.8 0.5 0.8 0.5 0.8 0.8 1.0 0.5 0.8 1.0 0.8 0.5 0.5 0.5
#> R003 0.5 0.5 0.8 1.0 0.5 0.8 0.8 0.8 1.0 1.2 0.5 0.8 0.5 0.8 0.2 0.8 0.5 0.5
#>       N1  N2  N3  N4  N5 N6  O1  O2  O3  O4  O5  O6
#> R001 0.8 0.8 0.8 0.5 0.8  1 0.5 1.0 0.2 0.8 1.0 0.8
#> R002 1.2 1.0 1.2 1.2 1.0  1 1.2 0.8 1.0 1.2 0.8 1.2
#> R003 1.0 1.0 0.8 0.5 0.5  1 0.2 1.0 0.2 0.5 0.2 0.5

Centering and scaling

scale() z-scores the active dimension (mean 0, sd 1), and center() only subtracts the mean.

Standardising items

Scaling columns puts every item on the same footing, whatever its mean and spread:

tm_z <- tm_scored |>
  activate(columns) |>
  scale()

round(colMeans(tm_z$matrix)[1:6], 3)
#> E1 E2 E3 E4 E5 E6 
#>  0  0  0  0  0  0
round(apply(tm_z$matrix, 2, sd)[1:6], 3)
#> E1 E2 E3 E4 E5 E6 
#>  1  1  1  1  1  1

scale() takes the usual center and scale arguments, so scale(center = TRUE, scale = FALSE) is equivalent to center().

Removing response style

Some people agree with nearly everything, others are more reserved. In psychology this is called acquiescence, and a simple way to reduce it is to center each respondent’s answers on their own mean (“ipsatising”). This is done on the raw answers: because positively and reverse-keyed items partly cancel out, a respondent’s raw mean mostly reflects their general tendency to agree.

tm_ipsative <- tm |>
  activate(rows) |>
  center()

round(rowMeans(tm_ipsative$matrix)[1:5], 3)
#> R001 R002 R003 R004 R005 
#>    0    0    0    0    0

Clipping

clip_values() caps values at a minimum and/or maximum. A typical use is to limit extreme z-scores before drawing a heatmap, so that a few outliers do not use up the whole colour scale:

range(tm_z$matrix)
#> [1] -2.566795  2.145061

tm_clipped <- tm_z |>
  activate(matrix) |>
  clip_values(min = -2, max = 2)

range(tm_clipped$matrix)
#> [1] -2  2

Log transformation

log_transform() is a convenience for skewed, count-like data such as gene expression or word counts, where a log2(x + 1) transformation is standard. It is not useful for Likert answers, so here is a small made-up count matrix:

counts <- tidymatrix(
  matrix(c(0, 3, 12, 150, 7, 0, 1024, 45, 2), nrow = 3),
  data.frame(gene = c("g1", "g2", "g3")),
  data.frame(sample = c("s1", "s2", "s3"))
)

counts |>
  activate(matrix) |>
  log_transform(base = 2, offset = 1) |>
  pull_active() |>
  round(2)
#>      [,1] [,2]  [,3]
#> [1,]  0.0 7.24 10.00
#> [2,]  2.0 3.00  5.52
#> [3,]  3.7 0.00  1.58

Transposing

t() swaps rows and columns, including the metadata. If rows were active, columns are active afterwards, so the active metadata stays the same.

tm_t <- t(tm_scored |> activate(rows))

dim(tm_scored$matrix)
#> [1] 388  30
dim(tm_t$matrix)
#> [1]  30 388

active(tm_t)
#> [1] "columns"
names(tm_t$row_data)
#> [1] "item_id"   "trait"     "reversed"  "item_text" "position"

Transposing is useful when a function only works along one direction, or when a package expects samples in rows rather than columns.

Getting the matrix out

pull_active() returns the active component: the matrix when the matrix is active, otherwise the active metadata.

m <- tm_scored |>
  activate(matrix) |>
  pull_active()

class(m)
#> [1] "matrix" "array"
dim(m)
#> [1] 388  30

Matrix operations and stored analyses

Every operation that changes the values makes stored analysis results (principal components, clusterings, …) out of date, so they are removed with a warning. The metadata columns they produced are kept. Transform first, analyse afterwards:

tm_pca <- tm_scored |>
  activate(columns) |>
  compute_prcomp(n_components = 2)

list_analyses(tm_pca)
#> [1] "column_pca"

tm_pca <- tm_pca |>
  activate(columns) |>
  scale()
#> Warning: Removed 1 stored analysis object(s) due to scale: column_pca
#> Metadata columns are preserved.

list_analyses(tm_pca)
#> character(0)

See also