dplyr verbs change which rows and columns a tidymatrix has. Matrix operations change the values. As elsewhere in the package, the active component decides the direction: with rows active an operation works row by row, with columns active column by column, and with the matrix active on all values at once.
library(tidymatrix)
library(dplyr, warn.conflicts = FALSE)
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)We use the big5 personality survey: 400 respondents
answering 30 items on a 1–5 scale (see ?big5).
Summary statistics into metadata
add_stats() computes statistics along the active
dimension and stores them as metadata columns. Per-respondent statistics
are a standard data quality check in surveys: someone who gives the same
answer to every question has a standard deviation close to zero.
tm <- tm |>
activate(rows) |>
add_stats(mean, sd, .names = c("resp_mean", "resp_sd"))
tm |>
activate(rows) |>
arrange(resp_sd) |>
select(respondent_id, resp_mean, resp_sd, completion_min)
#> # A tidymatrix: 400 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 400 rows x 4 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#> respondent_id resp_mean resp_sd completion_min
#> 1 R074 3.066667 0.3651484 2.0
#> 2 R307 3.900000 0.4025779 3.2
#> 3 R392 4.900000 0.4025779 1.5
#> 4 R170 3.966667 0.4138410 2.5
#> 5 R219 3.900000 0.5477226 2.2
#> 6 R352 4.866667 0.5713465 2.6The respondents with the least variable answers also finished the survey in a couple of minutes. Sorting by completion time shows the other kind of careless respondent — those who answered at random, and so have a normal-looking standard deviation:
tm |>
activate(rows) |>
arrange(completion_min) |>
select(respondent_id, resp_sd, completion_min)
#> # A tidymatrix: 400 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 400 rows x 3 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#> respondent_id resp_sd completion_min
#> 1 R392 0.4025779 1.5
#> 2 R282 1.3746473 1.7
#> 3 R393 1.2793677 1.7
#> 4 R074 0.3651484 2.0
#> 5 R368 1.4762487 2.0
#> 6 R219 0.5477226 2.2We remove both groups and move on with the clean data:
Per-item statistics work the same way with columns active. Custom
functions can be passed as a named list with .fns:
tm |>
activate(columns) |>
add_stats(.fns = list(
item_mean = mean,
item_sd = sd,
pct_agree = \(x) mean(x >= 4)
)) |>
select(item_id, item_text, item_mean, item_sd, pct_agree)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (columns):
#> item_id item_text item_mean item_sd pct_agree
#> 1 E1 I feel comfortable around people. 2.951031 1.145793 0.3350515
#> 2 E2 I start conversations. 3.087629 1.117191 0.3762887
#> 3 E3 I enjoy being part of a lively crowd. 3.373711 1.123868 0.4664948
#> 4 E4 I make friends easily. 3.257732 1.163869 0.4355670
#> 5 E5 I keep in the background. 2.935567 1.191474 0.3221649
#> 6 E6 I have little to say to strangers. 3.298969 1.096529 0.4432990Custom transformations with transform_matrix()
transform_matrix() applies any function to the matrix.
What the function receives depends on the active component:
| Active |
fn receives |
Example |
|---|---|---|
matrix |
the whole matrix |
log, \(m) m / 4
|
rows |
one row at a time | per-respondent ranks |
columns |
one column at a time | per-item standardisation |
The function must return values of the same size, and the row and
column names are preserved. Extra arguments are passed on to the
function; with rows active they are evaluated in the column metadata,
with columns active in the row metadata (use .env$x for a
variable x that clashes with a metadata column).
Reverse-scoring items
A questionnaire usually measures each trait with several statements, and the answers are later combined, for example averaged, into one score per trait. Some of the statements are deliberately worded in the opposite direction. These are called reverse-keyed items. They keep people who tend to agree with everything (or tick the same box throughout) from automatically ending up with high scores.
In big5, two of the six items per trait are
reverse-keyed. For example:
| Trait | Normally keyed | Reverse-keyed |
|---|---|---|
| Extraversion | “I start conversations.” | “I keep in the background.” |
| Conscientiousness | “I finish what I start.” | “I put off important tasks.” |
| Neuroticism | “I worry about things.” | “I stay calm under pressure.” |
Answering 5 (“strongly agree”) to “I keep in the background”
indicates low extraversion. Averaged as they are, such answers
would cancel out the normally keyed ones. So before the items are
combined, the answers to reverse-keyed items are flipped: 5 becomes 1, 4
becomes 2, 3 stays 3. On a 1–5 scale that is simply 6 - x.
Afterwards a high value always means a high trait level. This step is
known as reverse-scoring.
Which columns to flip is stored in the column metadata. With rows
active the function gets one respondent’s answers, one value per item.
Extra arguments to transform_matrix() are passed on to the
function, and they are evaluated with the column metadata as a data mask
— just like the arguments of mutate(). So
flip = reversed hands the function the
reversed column, which lines up element by element with the
answers:
tm_scored <- tm |>
activate(rows) |>
transform_matrix(\(x, flip) ifelse(flip, 6L - x, x), flip = reversed)
# E4 is positively keyed, E5 is reverse-keyed
tm$matrix[1:5, c("E4", "E5")]
#> E4 E5
#> R001 3 3
#> R002 2 3
#> R003 4 4
#> R004 3 3
#> R005 1 5
tm_scored$matrix[1:5, c("E4", "E5")]
#> E4 E5
#> R001 3 3
#> R002 2 3
#> R003 4 2
#> R004 3 3
#> R005 1 1After scoring, items of the same trait correlate positively:
Whole-matrix transformations
With the matrix active, the function gets the full matrix. Rescaling the 1–5 answers to a 0–100 scale:
tm_scored |>
activate(matrix) |>
transform_matrix(\(m) (m - 1) / 4 * 100) |>
pull_active() |>
head(3)
#> E1 E2 E3 E4 E5 E6 A1 A2 A3 A4 A5 A6 C1 C2 C3 C4 C5 C6 N1 N2 N3 N4 N5
#> R001 75 50 100 50 50 25 50 100 50 75 25 50 0 25 0 25 25 0 50 50 50 25 50
#> R002 25 75 25 25 50 25 50 25 50 50 75 25 50 75 50 25 25 25 100 75 100 100 75
#> R003 25 25 50 75 25 50 50 50 75 100 25 50 25 50 0 50 25 25 75 75 50 25 25
#> N6 O1 O2 O3 O4 O5 O6
#> R001 75 25 75 0 50 75 50
#> R002 75 100 50 75 100 50 100
#> R003 75 0 75 0 25 0 25Column- and row-wise transformations
With columns active the function is applied to each item separately. For example, percentile ranks within each item:
tm_scored |>
activate(columns) |>
transform_matrix(\(x) rank(x) / length(x)) |>
activate(matrix) |>
pull_active() |>
head(3) |>
round(2)
#> E1 E2 E3 E4 E5 E6 A1 A2 A3 A4 A5 A6 C1 C2 C3
#> R001 0.79 0.47 0.91 0.41 0.48 0.3 0.43 0.87 0.41 0.57 0.30 0.30 0.03 0.31 0.10
#> R002 0.24 0.76 0.14 0.17 0.48 0.3 0.43 0.10 0.41 0.27 0.86 0.11 0.37 0.84 0.61
#> R003 0.24 0.20 0.38 0.70 0.22 0.6 0.43 0.31 0.70 0.87 0.30 0.30 0.15 0.60 0.10
#> C4 C5 C6 N1 N2 N3 N4 N5 N6 O1 O2 O3 O4 O5 O6
#> R001 0.19 0.4 0.08 0.45 0.52 0.48 0.20 0.52 0.57 0.19 0.60 0.08 0.34 0.77 0.54
#> R002 0.19 0.4 0.29 0.93 0.80 0.93 0.93 0.79 0.57 0.92 0.32 0.81 0.90 0.50 0.95
#> R003 0.46 0.4 0.29 0.72 0.80 0.48 0.20 0.24 0.57 0.05 0.60 0.08 0.13 0.05 0.26Extra arguments are passed on to the function:
tm_scored |>
activate(columns) |>
transform_matrix(\(x) x / 4) |>
activate(matrix) |>
transform_matrix(round, digits = 1) |>
pull_active() |>
head(3)
#> E1 E2 E3 E4 E5 E6 A1 A2 A3 A4 A5 A6 C1 C2 C3 C4 C5 C6
#> R001 1.0 0.8 1.2 0.8 0.8 0.5 0.8 1.2 0.8 1.0 0.5 0.8 0.2 0.5 0.2 0.5 0.5 0.2
#> R002 0.5 1.0 0.5 0.5 0.8 0.5 0.8 0.5 0.8 0.8 1.0 0.5 0.8 1.0 0.8 0.5 0.5 0.5
#> R003 0.5 0.5 0.8 1.0 0.5 0.8 0.8 0.8 1.0 1.2 0.5 0.8 0.5 0.8 0.2 0.8 0.5 0.5
#> N1 N2 N3 N4 N5 N6 O1 O2 O3 O4 O5 O6
#> R001 0.8 0.8 0.8 0.5 0.8 1 0.5 1.0 0.2 0.8 1.0 0.8
#> R002 1.2 1.0 1.2 1.2 1.0 1 1.2 0.8 1.0 1.2 0.8 1.2
#> R003 1.0 1.0 0.8 0.5 0.5 1 0.2 1.0 0.2 0.5 0.2 0.5Centering and scaling
scale() z-scores the active dimension (mean 0, sd 1),
and center() only subtracts the mean.
Standardising items
Scaling columns puts every item on the same footing, whatever its mean and spread:
tm_z <- tm_scored |>
activate(columns) |>
scale()
round(colMeans(tm_z$matrix)[1:6], 3)
#> E1 E2 E3 E4 E5 E6
#> 0 0 0 0 0 0
round(apply(tm_z$matrix, 2, sd)[1:6], 3)
#> E1 E2 E3 E4 E5 E6
#> 1 1 1 1 1 1scale() takes the usual center and
scale arguments, so
scale(center = TRUE, scale = FALSE) is equivalent to
center().
Removing response style
Some people agree with nearly everything, others are more reserved. In psychology this is called acquiescence, and a simple way to reduce it is to center each respondent’s answers on their own mean (“ipsatising”). This is done on the raw answers: because positively and reverse-keyed items partly cancel out, a respondent’s raw mean mostly reflects their general tendency to agree.
Clipping
clip_values() caps values at a minimum and/or maximum. A
typical use is to limit extreme z-scores before drawing a heatmap, so
that a few outliers do not use up the whole colour scale:
range(tm_z$matrix)
#> [1] -2.566795 2.145061
tm_clipped <- tm_z |>
activate(matrix) |>
clip_values(min = -2, max = 2)
range(tm_clipped$matrix)
#> [1] -2 2Log transformation
log_transform() is a convenience for skewed, count-like
data such as gene expression or word counts, where a
log2(x + 1) transformation is standard. It is not useful
for Likert answers, so here is a small made-up count matrix:
counts <- tidymatrix(
matrix(c(0, 3, 12, 150, 7, 0, 1024, 45, 2), nrow = 3),
data.frame(gene = c("g1", "g2", "g3")),
data.frame(sample = c("s1", "s2", "s3"))
)
counts |>
activate(matrix) |>
log_transform(base = 2, offset = 1) |>
pull_active() |>
round(2)
#> [,1] [,2] [,3]
#> [1,] 0.0 7.24 10.00
#> [2,] 2.0 3.00 5.52
#> [3,] 3.7 0.00 1.58Transposing
t() swaps rows and columns, including the metadata. If
rows were active, columns are active afterwards, so the active
metadata stays the same.
tm_t <- t(tm_scored |> activate(rows))
dim(tm_scored$matrix)
#> [1] 388 30
dim(tm_t$matrix)
#> [1] 30 388
active(tm_t)
#> [1] "columns"
names(tm_t$row_data)
#> [1] "item_id" "trait" "reversed" "item_text" "position"Transposing is useful when a function only works along one direction, or when a package expects samples in rows rather than columns.
Getting the matrix out
pull_active() returns the active component: the matrix
when the matrix is active, otherwise the active metadata.
m <- tm_scored |>
activate(matrix) |>
pull_active()
class(m)
#> [1] "matrix" "array"
dim(m)
#> [1] 388 30Matrix operations and stored analyses
Every operation that changes the values makes stored analysis results (principal components, clusterings, …) out of date, so they are removed with a warning. The metadata columns they produced are kept. Transform first, analyse afterwards:
tm_pca <- tm_scored |>
activate(columns) |>
compute_prcomp(n_components = 2)
list_analyses(tm_pca)
#> [1] "column_pca"
tm_pca <- tm_pca |>
activate(columns) |>
scale()
#> Warning: Removed 1 stored analysis object(s) due to scale: column_pca
#> Metadata columns are preserved.
list_analyses(tm_pca)
#> character(0)See also
- PCA and clustering builds on the scored matrix.
- Plotting and exporting shows how to use scaled and clipped data in a heatmap.