This case study shows the intended gtstats workflow. A maternity team wants to describe participants by low-birth-weight outcome, understand the continuous variables, and produce a Table 1. The example uses the labelled birth-weight teaching dataset included with gtstats.
The built-in data keep the original code names but supply readable
variable labels and factors. antenatal_visits is explicitly
ordered; numeric clinical codes are not silently treated as ordinal
merely because their values are ordered.
Variable | Type | Complete | Unique | Overview | Range / levels |
|---|---|---|---|---|---|
Birth-weight outcome [low] | binary | 189/189 (100.0%) | 2 | Normal birth weight 130 (68.8%); Low birth weight 59 (31.2%) | Normal birth weight, Low birth weight |
Maternal age (years) [age] | continuous | 189/189 (100.0%) | 24 | Mean 23.24 (SD 5.30); median 23.00 | 14.00 to 45.00 |
Maternal weight (lb) [lwt] | continuous | 189/189 (100.0%) | 75 | Mean 129.81 (SD 30.58); median 121.00 | 80.00 to 250.00 |
Maternal race [race] | categorical | 189/189 (100.0%) | 3 | White 96 (50.8%); Other 67 (35.4%); Black 26 (13.8%) | White, Black, Other |
Smoking during pregnancy [smoke] | binary | 189/189 (100.0%) | 2 | No 115 (60.8%); Yes 74 (39.2%) | No, Yes |
Previous premature labours [ptl] | categorical* | 189/189 (100.0%) | 4 | 0 159 (84.1%); 1 24 (12.7%); 2 5 ( 2.6%) | 0, 1, 2, 3 |
Hypertension [ht] | binary | 189/189 (100.0%) | 2 | No 177 (93.7%); Yes 12 ( 6.3%) | No, Yes |
Uterine irritability [ui] | binary | 189/189 (100.0%) | 2 | No 161 (85.2%); Yes 28 (14.8%) | No, Yes |
First-trimester visits [ftv] | continuous | 189/189 (100.0%) | 6 | Mean 0.79 (SD 1.06); median 0.00 | 0.00 to 6.00 |
Birth weight (g) [bwt] | continuous | 189/189 (100.0%) | 131 | Mean 2944.59 (SD 729.21); median 2977.00 | 709.00 to 4990.00 |
Previous premature labour [previous_preterm] | binary | 189/189 (100.0%) | 2 | No 159 (84.1%); Yes 30 (15.9%) | No, Yes |
First-trimester visits [antenatal_visits] | ordinal | 189/189 (100.0%) | 3 | None 100 (52.9%); One 47 (24.9%); Two or more 42 (22.2%) | None, One, Two or more |
One row is shown per selected variable. | |||||
Potential data-quality findings and interpretation prompts are stored in `$issues`. | |||||
* Possible ordinal or count-coded variable; confirm the intended meaning and order from the data dictionary or clinical context. | |||||
overview$issues
#> # A tibble: 2 × 5
#> variable label issue why_flagged suggested_check
#> <chr> <chr> <chr> <chr> <chr>
#> 1 ptl Previous premature labours Sparse catego… At least o… Confirm coding…
#> 2 ftv First-trimester visits Low-cardinali… Only 6 dis… Confirm whethe…describe_data() is the first pass. It reports type,
completeness, levels or range, and a short type-specific overview. It
does not decide a statistical test.
distribution <- assess_distribution(
birthwt_data,
vars = c(age, lwt),
by = low
)
to_flextable(distribution)Variable | Group | n | Skewness | Shape | Suggested presentation | Shapiro p |
|---|---|---|---|---|---|---|
Maternal age (years) | Normal birth weight | 130 | 0.74 | Some right asymmetry | Review mean (SD) and median (IQR) | <0.001 |
Low birth weight | 59 | 0.29 | Little/no asymmetry | 0.521 | ||
Maternal weight (lb) | Normal birth weight | 130 | 1.42 | Marked right skew | Median (IQR) preferred | <0.001 |
Low birth weight | 59 | 1.06 | Marked right skew | <0.001 | ||
Shape categories use absolute sample skewness: little/no asymmetry < 0.50; some asymmetry 0.50 to < 1.00; marked skew >= 1.00. They are descriptive guidance, not formal classifications. | ||||||
Shapiro-Wilk is sensitive to sample size. Interpretation should consider skewness, graphical assessment, sample size and subject-matter knowledge. | ||||||
Suggested summaries are intended for descriptive reporting only and should not be used alone to determine inferential statistical methods. | ||||||
For grouped data, the suggested presentation applies to all groups of each variable. | ||||||
distribution$recommendations
#> # A tibble: 2 × 5
#> variable label overall_recommendation reason review
#> <chr> <chr> <chr> <chr> <chr>
#> 1 age Maternal age (years) Review mean (SD) and median (IQR) Some a… Inspe…
#> 2 lwt Maternal weight (lb) Median (IQR) preferred Marked… Inspe…The first table retains group-specific diagnostics. The
recommendation table gives one reporting approach per variable, so the
same summary can be used across Table 1 groups. Shapiro-Wilk is
supporting information, not a rule for choosing an inferential test. Use
plots = TRUE when a histogram, density, Q-Q plot, and
boxplot are useful.
assess_variance() provides the related spread
diagnostic. It reports group SDs, variances, and ratios without using a
variance test to gatekeep Welch methods.
Variable | Normal birth weight | Low birth weight | Observed SD ratio | Observed variance ratio | Levene p | Interpretation |
|---|---|---|---|---|---|---|
Maternal age (years) | n = 130; SD = 5.58; variance = 31.19 | n = 59; SD = 4.51; variance = 20.35 | 1.24 | 1.53 | 0.102 | Levene found no clear evidence of different spreads; this does not prove equal variances. |
Maternal weight (lb) | n = 130; SD = 31.72; variance = 1006.41 | n = 59; SD = 26.56; variance = 705.40 | 1.19 | 1.43 | 0.476 | Levene found no clear evidence of different spreads; this does not prove equal variances. |
SD and variance ratios are the largest group value divided by the smallest group value. They describe observed spread; they are not pass/fail tests. | ||||||
For independent groups, Welch t-tests and Welch ANOVA do not require equal variances. `assess_variance()` does not select an inferential test and does not assess pairing or repeated-measures sphericity. | ||||||
The displayed Levene test is median-centred (Brown-Forsythe). It is supporting information only: its p-value neither proves equal variances nor selects an inferential test. | ||||||
Interpret spread alongside sample size, distributional shape, outliers, missingness, and the study design. | ||||||
table_one <- summary_table(
birthwt_data,
by = low,
include = c(age, lwt, race, smoke, ht, ui,
previous_preterm, antenatal_visits),
overall = "first"
) |>
add_p()
to_flextable(table_one)Characteristic | Overall | Normal birth weight | Low birth weight | p-value |
|---|---|---|---|---|
Maternal age (years) | 23.2 (5.3) | 23.7 (5.6) | 22.3 (4.5) | 0.078ᵃ |
Maternal weight (lb) | 121.0 (110.0–140.0) | 123.5 (113.0–147.0) | 120.0 (104.0–130.0) | 0.013ᵇ |
Maternal race | 0.082ᶜ | |||
White | 96 (50.8%) | 73 (56.2%) | 23 (39.0%) | |
Black | 26 (13.8%) | 15 (11.5%) | 11 (18.6%) | |
Other | 67 (35.4%) | 42 (32.3%) | 25 (42.4%) | |
Smoking during pregnancy | 0.040ᶜ | |||
No | 115 (60.8%) | 86 (66.2%) | 29 (49.2%) | |
Yes | 74 (39.2%) | 44 (33.8%) | 30 (50.8%) | |
Hypertension | 0.052ᵈ | |||
No | 177 (93.7%) | 125 (96.2%) | 52 (88.1%) | |
Yes | 12 (6.3%) | 5 (3.8%) | 7 (11.9%) | |
Uterine irritability | 0.035ᶜ | |||
No | 161 (85.2%) | 116 (89.2%) | 45 (76.3%) | |
Yes | 28 (14.8%) | 14 (10.8%) | 14 (23.7%) | |
Previous premature labour | <0.001ᶜ | |||
No | 159 (84.1%) | 118 (90.8%) | 41 (69.5%) | |
Yes | 30 (15.9%) | 12 (9.2%) | 18 (30.5%) | |
First-trimester visits | 0.281ᶜ | |||
None | 100 (52.9%) | 64 (49.2%) | 36 (61.0%) | |
One | 47 (24.9%) | 36 (27.7%) | 11 (18.6%) | |
Two or more | 42 (22.2%) | 30 (23.1%) | 12 (20.3%) | |
Continuous data: Maternal age (years): mean (SD); Maternal weight (lb): median (IQR). Categorical data are n (%). | ||||
ᵃ Welch t-test; ᵇ Wilcoxon rank-sum test; ᶜ Chi-square test; ᵈ Fisher's exact test | ||||
This is already a publication-ready flextable. Use
customise_table() to finish it for Word; use
to_gt() only when a gt object is specifically needed.
weight_comparison <- compare_groups(
birthwt_data,
variable = lwt,
group = low
)
to_flextable(weight_comparison)Variable | Normal birth weight | Low birth weight | p-value |
|---|---|---|---|
Maternal weight (lb) | 123.50 (113.00–147.00) | 120.00 (104.00–130.00) | 0.01 |
The analysis assumes independent observations; study design must confirm this. | |||
Automatic selection used distribution guidance or design/cell-count rules. Rule applied: Two-group continuous outcome: marked skewness flagged in at least one group; selected Wilcoxon rank-sum test. | |||
Wilcoxon rank-sum is rank based; similar distribution shapes are needed for a median-shift interpretation. | |||
# The exact automatic rule and observed inputs are retained for review.
weight_comparison$method$selection_rule
#> [1] "Two-group continuous outcome: marked skewness flagged in at least one group; selected Wilcoxon rank-sum test."
weight_comparison$method$selection_inputs
#> $distribution_guidance
#> [1] "Skewed; Skewed"
#>
#> $skewness_flagged
#> [1] TRUE
#>
#> $groups
#> [1] 2
#>
#> $paired
#> [1] FALSE
#>
#> $var_equal
#> [1] FALSE
#>
#> $variance_assumption_source
#> [1] "User-specified; not inferred from a variance hypothesis test"
#>
#> $observed_group_spread
#> $observed_group_spread$group_spread
#> # A tibble: 2 × 4
#> group n sd variance
#> <chr> <int> <dbl> <dbl>
#> 1 Normal birth weight 130 31.7 1006.
#> 2 Low birth weight 59 26.6 705.
#>
#> $observed_group_spread$sd_ratio
#> [1] 1.194461
#>
#> $observed_group_spread$variance_ratio
#> [1] 1.426737
assumptions_stats(weight_comparison)| Checks before reporting | |||
| Variable | Check before reporting | Action | Details |
|---|---|---|---|
| Maternal weight (lb) | Independent observations | Confirm from study design | Confirm from the study design that each observation contributes independently. |
| Maternal weight (lb) | Comparable distribution shapes | Confirm from study design | Required when interpreting the rank-based result specifically as a location or median shift. |
| Diagnostics | |||||
| Variable | Check | Result | Observed value | Reference | Interpretation |
|---|---|---|---|---|---|
| Maternal weight (lb) | Comparison design | independent | Independent observations | Defined by the study design | Independent comparison; confirm independence from the study design. |
| Maternal weight (lb) | Variance assumption | welch default | var_equal = FALSE | User-specified analytical assumption | Welch is the conservative default and does not require equal variances. No variance hypothesis test was used. |
| Maternal weight (lb) | Automatic test selection | Wilcoxon rank-sum test | Skewed; Skewed | No marked group-level skewness flag for parametric default | Two-group continuous outcome: marked skewness flagged in at least one group; selected Wilcoxon rank-sum test. |
| Maternal weight (lb) | Distribution guidance | rank based recommended | Skewed; Skewed | Marked absolute skewness guidance; Shapiro-Wilk is supporting information only | Assessment applied within each group. |
| Maternal weight (lb) | Observed group spread | descriptive context | Normal birth weight (n = 130): SD 31.72; variance 1006.41; Low birth weight (n = 59): SD 26.56; variance 705.40; SD ratio = 1.19; variance ratio = 1.43 | Descriptive diagnostic; no pass/fail threshold | Observed spread is descriptive only; no Levene, Bartlett, or F test was performed. `var_equal = FALSE` retains Welch methods as the conservative parametric auto default. |
| Denominator audit | ||||||||
| Variable | Level | Group | Eligible observations | Used in analysis | Missing / excluded | Numerator | Denominator | Rule |
|---|---|---|---|---|---|---|---|---|
| lwt | NA | low = Normal birth weight | 130 | 130 | 0 | NA | 130 | Non-missing outcome observations within group |
| lwt | NA | low = Low birth weight | 59 | 59 | 0 | NA | 59 | Non-missing outcome observations within group |
compare_groups() answers one question at a time. The
publication table is the result; the helper functions expose the
assumptions, automatic checks, and analysis denominators for review. Use
effect_size() separately when the magnitude of the
difference is the main question.