T test

Goran Kardum

Department of Psychology

Prerequisites

  • Concept and features of normal distribution
  • Descriptive statistics (arithmetic mean and standard deviations)
  • Standard error of mean and deviations
  • Standard error of difference between two arithmetic means
  • Distribution of a large number of samples

Standard error and C.I.

  • For the correct use of the t-test and other statistical tests, a detailed and very good understanding of the relationship between the sample and the population is necessary, i.e. understanding of the concept of error of measurement the error related to the individual arithmetic mean and standard deviation

  • … and… also errors related to the difference between two arithmetic means, i.e. standard error of the difference between two arithmetic means.

Population and sample

  • Population consists of all individuals or entities that have some common property or properties.

  • A sample is a part of the population that can have different characteristics and a certain degree of representativeness.

  • A sample is always an estimate of the population and a certain error is always associated with it.

Population and sample

  • The arithmetic mean of the sample is always an estimate and deviates to a certain extent from the true/real (population, denoted by \(\mu\)) arithmetic mean.

  • When taking random samples, we should also take into account random variations of the subject matter properties of the object of measurement among the samples. Of course, the greater the variability of dependent variable/subject matter, the greater the variability between samples.

Distribution of a large number of samples

  • The standard deviation of the arithmetic means of the samples is smaller the larger the sample is.

  • The larger the sample, the closer we are to estimating the true arithmetic mean, i.e. the arithmetic mean of the sample is closer to the true arithmetic mean.

  • The distribution of arithmetic means will follow the form of a normal distribution, even when the object of measurement from the population is not normally distributed.

  • We call this phenomenon the central limit theorem.

Distribution of samples

  • The distribution of the arithmetic means of the samples approaches to normal distribution as N of samples increases.

  • The larger the sample and the smaller the variability of the occurrence in the population, the more accurate the estimation of the population parameters from the sample.

Se arithmetic mean

  • A statistical indicator that is very important for this discussion: the standard error of the arithmetic mean

  • Estimate the deviation of large number of the samples arithmetic means from the true arithmetic mean.

  • The calculation formula takes into account the following: variability and sample size.

\[\begin{equation}\label{pogreška aritmetičke sredine} SE_{\overline{x}}=\frac{SD}{\sqrt{N}} \end{equation}\]

Se two arithmetic means

  • In the following equation, there is a formula for calculating the standard error between two arithmetic means.

  • It is basically the sum of errors that are “attached” to one and the other arithmetic mean.

\[\begin{equation}\label{pogreška razlika između aritmetičkih sredina} Se_{M_{1}-M_{2}}=\sqrt{SD^{2}_{M_{1}}+SD^{2}_{M_{2}}} \end{equation}\]

T test - general consideration

  • Gosset (William Sealy Gosset, 1876-1937) published under the name of Student and developed Student’s t-distribution (Boland, 1984)

  • Head Experimental Brewer at Guinness

  • A pioneer of modern statistics

  • Experimental design and analysis with an economic approach to the logic of uncertainty

  • Published under the pen name Student and developed most famously Student’s t-distribution, originally called Student’s “z”

T test for independent samples

  • to compare the means of two independent samples on a given variable.

  • the standard Student’s t-test, which assumes that the variance of the two groups are equal

  • the Welch’s t-test, which is less restrictive compared to the original Student’s test: do not assume that the variance is the same in the two groups, which results in the fractional degrees of freedom.

  • if you wanted to compare the average number of correct anwers on knowledge test of 50 randomly selected men to that of 50 randomly selected women, you would conduct an independent samples t test.

T value - independent t-test

\[\begin{equation} t = \frac{\bar X1 - \bar X2}{\sqrt{Se1 + Se2}} \end{equation}\]

T - test

  • Parametric statistical test

  • To test difference between two samples

  • Dependent variable - on interval or ratio scale

  • Normal distribution

  • (over)used tool in all statistics

T test - general consideration

  • Number of cases or subjects (N>30)

  • Symmetrical Distribution

  • Equality of variances

  • Take care of outliers

T test

  • T - test (one sample): used to compare the mean of a population with a theoretical value (or hypothesis, results of some other research)

  • T - test (two independent sample, used to compare the mean of two independent samples)

  • T - test (two related, dependent sample, used to compare the means between two related groups of samples)

One-Sample t test

  • The one-sample t-test is a statistical test used to determine whether the mean of a single sample of data is significantly different from a known or hypothesized value. It is a common test used in inferential statistics to draw conclusions about a population based on a sample of data.

  • The basic idea behind the one-sample t-test is to compare the sample mean to a hypothesized population mean, and determine whether the difference between the two is large enough to be statistically significant. The test assumes that the data are normally distributed and that the population variance is unknown but can be estimated from the sample.

  • purpose of test: a one-sample t test is performed when you want to compare the mean from a single sample to a population mean.

Applications, One-Sample t test

Typical applications of the one-sample t-test include:

  • Researches may wish to know if their data differs from an external standard (from manual, handbooks)

  • Testing whether a sample mean is significantly different from a known value, such as the population mean or a hypothesized value based on theory or previous research.

  • Comparing the effectiveness of a new treatment or intervention to a known or established standard.

Applications, One-Sample t test

  • Testing whether a process or system is operating within acceptable limits by comparing the observed mean to a specified target value.

  • Evaluating the accuracy of a measuring instrument by comparing the mean of a sample of measurements to a known or accepted standard.

Assumptions of the one-sample t-test (Navarro, 2015)

There are a number of assumptions that need to be met in order for a one-sample t-test to be valid. Some of these are more important than others;

  • Independence: In rough terms, independence means each observation in the sample does not ‘depend on’ the others. The key thing to know now is why this assumption matters: if the data are not independent the p-values generated by the one-sample t-test will be unreliable.

  • the p-values will be too small when the non-independence assumption is broken. That means we risk the false conclusion that a difference is statistically significant, when in reality, it is not

Assumptions of the one-sample t-test

  • Measurement scale: The variable being analysed should be measured on an interval or ratio scale, i.e. it should be a numeric variable of some kind. It doesn’t make much sense to apply a one-sample t-test to a variable that isn’t measured on one of these scales.

  • Normality: The one-sample t-test will only produce completely reliable p-values when the variable is normally distributed in the population. However, this assumption is less important than many people think. The t-test is robust to mild departures from normality when the sample size is small, and when the sample size is large the normality assumption hardly matters at all.

T value - one sample t-test

\[\begin{equation} t = \frac{\bar X - \mu}{s / \sqrt{n}} \end{equation}\]

Statistical hypotheses

  • Ho: m = μ
  • Ho: m ≤ μ
  • Ho: m ≥ μ

Example of one-sample t test

Chao, C.-N. (2017). An Examination on the Chinese Students’ Rationales to Receive their Higher Education in the U.S. World Journal of Education, 7(3), 41. https://doi.org/10.5430/wje.v7n3p41

  • Why those authors used one-sample t test?
  • Estimate value for testing hypotheses in this study is 3. Why?

Paired t-test

  • Calculate the difference (d) between each pair of value
  • Compute the mean (m) and the standard deviation (s) of d
  • Compare the average difference to 0. If there is any significant difference between the two pairs of samples, then the mean of d (m) is expected to be far from 0.

T value - paired t-test

\[\begin{equation} t = \frac{m}{s / \sqrt{n}} \end{equation}\]

Checking Assumptions

  • measurement scale

  • number of cases

  • normality of distribution

  • symmetry of distribution

  • homogeneity of variance

Degrees of freedom

  • Degrees of freedom (abbreviated d.f. or df) are closely related to the idea of sample size. The greater the degrees of freedom associated with a test, the more likely it is to detect an effect if it’s present. To calculate the degrees of freedom, we start with the sample size and then we reduce this number by one for every quantity (e.g. a mean) we had to calculate to construct the test.

  • Calculating degrees of freedom for a one-sample t-test is easy. The degrees of freedom are just n-1, where n is the number of observations in the sample. We lose one degree of freedom because we have to calculate one sample mean to construct the test.

Effect size

  • power analysis

  • sample size

  • number of cases/subjects (total sample size calculation)

  • type of statistical tools/tests

Kolmogorov - Smirnov test

  • used to test whether or not or not a sample comes from a certain distribution.

  • one or two samples

  • produce a D along with a corresponding p - value

  • ks.test() function in R

Shapiro-Wilk test

  • test of normality

  • whether or not a sample comes from a normal distribution

  • assumption used in many statistical tests t-test, regression analysis, ANOVA

  • shapiro.test() function in R

  • produces a W along with a corresponding p-value. If the p-value is less than α =.05, there is sufficient evidence that the sample does not come from a population that is normally distributed

Homogeneity of variance

  • the variance across our samples should be roughly equal

  • Bartlett test for parametric data (bartlett.test()) (Williams, n.d.)

  • Fligner-Killeen test for non-parametric data (fligner.test()).

Confidence in sample mean

  • Ratio of mean and standard deviation: mean/sd

  • Number that tell us about confidence in the mean!

  • After dividing these two numbers. What happens?

  • We could get smaller number, 1 and big number

  • Big number: more confident that the mean represents all of the numbers, respondents

Important ratio

  • ratio of measure of what we know dividing by measure of what we don’t know

  • In other words….

  • ratio of measure of effect / measure of error (Dancey, 2020)

ASA statements about p-values

P-values intepretations are widely misused (Wasserstein & Lazar, 2016)

  • P-values can indicate how incompatible the data are with a specified statistical model.

  • P-values do not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone.

  • Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold.

… from previous

  • Proper inference requires full reporting and transparency.

  • A p-value, or statistical significance, does not measure the size of an effect or the importance of a result.

  • By itself, a p-value does not provide a good measure of evidence regarding a model or hypothesis.

rempsyc package (Thériault, 2023)

  • nice_t_test function

  • Easily compute t-test analyses, with effect sizes, and format in publication-ready format.

  • The 95% confidence interval is for the effect size, Cohen’s d, both provided by the effectsize package.

library(rempsyc)

t.tests <- nice_t_test(
  data = mtcars,
  response = c("mpg", "disp", "drat", "wt"),
  group = "am"
)

t.tests
  Dependent Variable         t       df            p         d   CI_lower
1                mpg -3.767123 18.33225 1.373638e-03 -1.477947 -2.2659732
2               disp  4.197727 29.25845 2.300413e-04  1.445221  0.6417836
3               drat -5.646088 27.19780 5.266742e-06 -2.003084 -2.8592770
4                 wt  5.493905 29.23352 6.272020e-06  1.892406  1.0300224
    CI_upper
1 -0.6705684
2  2.2295594
3 -1.1245499
4  2.7329219

Format table (APA style)

t_table <- nice_table(t.tests)
t_table

Dependent Variable

t

df

p

d

95% CI

mpg

-3.77

18.33

.001**

-1.48

[-2.27, -0.67]

disp

4.20

29.26

< .001***

1.45

[0.64, 2.23]

drat

-5.65

27.20

< .001***

-2.00

[-2.86, -1.12]

wt

5.49

29.23

< .001***

1.89

[1.03, 2.73]

nice_violin(
  data = ToothGrowth,
  group = "dose",
  response = "len",
  xlabels = c("Low", "Medium", "High"),
  comp1 = 1,
  comp2 = 3,
  has.d = TRUE,
  d.y = 30
)

nice_normality(
  data = iris,
  variable = "Sepal.Length",
  group = "Species",
  grid = FALSE,
  shapiro = TRUE,
  histogram = TRUE
)

Visually check outliers based on (e.g.) +/- 3 MAD (median absolute deviations) or SD (standard deviations).

plot_outliers(airquality,
  group = "Month",
  response = "Ozone"
)

Test your knowledge

  • Sleep researchers decide to test the impact of REM sleep deprivation on cognitive performance. Subjects are required to participate in two nights of testing. The task involves five rows of widgets slowly passing across the computer screen. Randomly placed on a one/five ratio are widgets missing a component that must be “fixed” by the subject. Number of missed widgets is recorded. Compute the appropriate t-test for the data provided below (number of errors). (webster.edu)

  • Control condition: 20,4,9,36,20,3,25,10,6,14

  • REM deprived: 26,15,8,44,26,13,38,24,17,29

Questions

  • Type of t-test?

  • What would be the null and alternate hypothesis in this study?

  • What probability level did you choose and why?

  • What is your critical t-value? Is there a significant difference between the two testing conditions?

  • Calculate power and d. If you have made an error, would it be a Type I or a Type II error?

Null hypothesis controversy

The Null Hypothesis Testing Controversy in Psychology Author(s): David H. Krantz

Source: Journal of the American Statistical Association, Vol. 94, No. 448 (Dec., 1999), pp. 1372-1381

Published by: Taylor & Francis, Ltd. on behalf of the American Statistical Association Stable URL: https://www.jstor.org/stable/2669949

Misreporting of statistical results

The (mis)reporting of statistical results in psychology journals Authors Marjan Bakker & Jelte M. Wicherts

Behavioral Research (2011) 43:666–678, DOI 10.3758/s13428-011-0089-5

Statistical assumptions and reproducibility

Zhang W, Yan S, Tian B and Fei D (2022) Statistical Assumptions and Reproducibility in Psychology: Data Mining Based on Open Science. Front. Psychol. 13:905977. doi: 10.3389/fpsyg.2022.905977

Psychologists and Welch t-test

Delacre, M., et al (2017). Why Psychologists Should by Default Use Welch’s t-test Instead of Student’s t-test. International Review of Social Psychology, 30(1), 92–101, DOI: https://doi.org/10.5334/irsp.82