Skip to content

Add graph_test_shortcut_gsDesign(): group sequential graphical testing backed by gsDesign - #109

Open
xidongdxi wants to merge 6 commits into
mainfrom
gsd-engine-injection
Open

xidongdxi wants to merge 6 commits into
mainfrom
gsd-engine-injection

Conversation

@xidongdxi

@xidongdxi xidongdxi commented Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR separates the group sequential calculations from the graphical calculations in graph_test_shortcut_gsd() and adds graph_test_shortcut_gsDesign(), which runs the same graphical group sequential procedure but obtains every group sequential quantity (boundaries and repeated p-values) from the gsDesign package, using gsDesign's own spending terms (sfu, sfupar, usTime). It comes with a dedicated vignette that reproduces the group sequential testing vignette with the new function.

gsDesign stays in Suggests and is only required by the new function, which stops with an informative message if it is missing. The behavior of graph_test_shortcut_gsd() is unchanged.

The PR is organized in four layers; the commits follow the same order.

1. Engine injection (behavior-neutral refactor)

The graphical loop in gsd_test() needs exactly two group sequential operations per hypothesis: a repeated p-value and a set of nominal boundaries. These are now supplied by an engine, a list of two closures repeated_p(p, info_frac, j) and nominal_bounds(alpha, info_frac, j) (see R/gsd_engine.R). gsd_engine_mvtnorm(spending_fn) wraps the existing repeated_p()/gs_boundaries(), and gsd_test(), gsd_test_values_details(), gsd_test_values_look_back() and gsd_boundary_table() take an engine instead of spending_fn. seq_p = cummin(rep_p) stays in the graphical layer, so it is engine-agnostic.

Neutrality was verified with the full existing test suite and by knitting both group sequential vignettes under a fixed seed against main: the rendered markdown is byte-identical.

2. gsDesign engine (gsd_engine_gsDesign(), internal)

Only gsDesign's public functions are used: gsDesign(), sequentialPValue(), and the sf* spending functions.

  • One design per hypothesis from its non-NA analyses: gsDesign(k = K_j, test.type = 1, alpha, sfu, sfupar, n.I = info_frac_j, maxn.IPlan = 1, usTime) — one-sided efficacy boundary, information given directly as fractions of the planned maximum, so the correlation is the one implied by info_frac. The design is rebuilt at whatever level the graph currently allocates (weight × alpha).
  • Repeated p-values come from sequentialPValue() after setting the Z-statistics of analyses 1..k-1 to -20 (gsDesign's own "no lower bound" value), so that only analysis k can trigger a rejection. This is exact, not an approximation. gsDesign has no repeatedPValue() function (checked in 3.9.0 and CRAN 3.11.0).
  • The spending time is always passed explicitly (usTime if given, else pmin(info_frac, 1)). sequentialPValue()'s default rescales it by the last observed information, which silently changes interim spending when a trial over-runs its plan (error 4e-4 vs 2e-7 with an explicit usTime).
  • Classical families "OF", "Pocock", "WT" are boundary shapes in gsDesign, not spending functions, and sequentialPValue() errors on them; their repeated p-value is found by a root-find of b_k(alpha) = Z_k on the design's upper bound (beta = (1 - alpha)/2 keeps gsDesign() valid near alpha ≈ 1; beta does not affect the bounds for test.type = 1). Wang-Tsiatis repeated p-values were cross-validated against rpact's exact WT critical values (≤ 1e-7).
  • Single-analysis hypotheses use the univariate closed form (gsDesign() requires k >= 2).
  • Edge cases: allocated level 0 / 1, extreme p-values, and O'Brien-Fleming spending that underflows to exactly 0 at early spending times (sequentialPValue() would error "Final spend must be > 0"; the search interval is raised to the first level with positive spend).
  • A user-defined spending function built on gsDesign::spendingFunction() is stored in the design (upper$sf) so that sequentialPValue() evaluates it rather than the template's linear spend.

3. graph_test_shortcut_gsDesign() (exported)

Same signature as graph_test_shortcut_gsd() with spending_fn replaced by sfu (spending function or one of "OF", "Pocock", "WT"; single value or list of m), sfupar (single value or list of m), and usTime (NULL, vector of length K, or m × K matrix with info_frac's NA pattern). Returns the same gsd_graph_report structure, with inputs$engine = "gsDesign"; the print method labels spending functions by gsDesign's own $name and shows usTime. Validation covers sfu/sfupar/usTime shapes and contents, warns that usTime is ignored for classical families, and rejects gsDesign::sfPoints() (it needs the full schedule at every call, which sequentialPValue() cannot provide; sfLinear() gives the same spending). The roxygen @details describes how gsDesign is used.

Mapping of the built-in spending functions: spending_of = sfLDOF, spending_pocock = sfLDPocock, spending_hsd(gamma) = sfHSD + sfupar = gamma, spending_linear = sfPower + sfupar = 1, spending_wt(delta) = "WT" + sfupar = delta, spending_with_time() = usTime.

Agreement with graph_test_shortcut_gsd(): gsDesign integrates deterministically, the mvtnorm engine uses randomized quasi-Monte Carlo, so results agree to about 1e-5 (repeated/sequential/adjusted p-values within 8e-6, boundaries within 1e-8) rather than bit-for-bit; rejection decisions, sequences, and test_values coincide in all test scenarios (Maurer–Bretz look_back F/T, oncology NA-padded schedules, per-hypothesis spending/info fractions/look_back, over-running information, spending time, single-analysis hypotheses).

4. Vignette

vignettes/group-sequential-testing-gsDesign.Rmd ("Graphical Approaches in Group Sequential Design: Combining graphicalMCP and gsDesign") mirrors group-sequential-testing.Rmd section by section with graph_test_shortcut_gsDesign(). The original vignette is retitled "Graphical Approaches in Group Sequential Design" so that the two read as a pair and neither suggests that gsDesign handles the graphical part; file names and pkgdown URLs are unchanged. Contents of the new vignette: gsDesign's spending functions and a correspondence table; boundaries from gsDesign(); an explicit demonstration of the sequentialPValue() + -20 mechanism (what the function does internally); both case studies with identical numbers; spending time via usTime; monitoring via usTime = planned, info_frac = actual, plus the equivalent over-run representation; gsDesign's spending library (incl. "WT", sfTruncated, and why sfPoints is unsupported); a user-defined spending function on gsDesign::spendingFunction(). Per our discussion it has no cross-check section against graph_test_shortcut_gsd() (a separate vignette later; the equivalence lives in the tests) and no rpact material. If gsDesign is not installed, a visible note is printed and knitting stops (knitr::knit_exit()), which also covers the inline code. The original vignette's "Using Spending Functions from Other Packages" section now points to the new function, and the article is added to _pkgdown.yml and NEWS.

Verification

  • devtools::test(): 1048 pass, 0 fail, 1 skip (lrstat ≥ 0.3.3 not installed locally); the original 572 tests are unchanged.
  • devtools::check(args = c("--as-cran", "--no-manual")): 0 errors, 0 warnings; the only notes are the standard "development version" NEWS heading and a local scratch directory that is not in the branch.
  • Both vignettes knit; the new one's numbers match the text of the original (e.g. analysis-1 boundary 0.000015, analysis-2 boundaries 0.0022/0.0040/0.0022/0.0060).

Suggested reading order for review

  1. R/gsd_engine.R — gsd_engine_mvtnorm() then gsd_engine_gsDesign() (the whole gsDesign interaction is here, ~120 lines).
  2. R/graph_test_shortcut_gsd.R — the engine plumbing (gsd_test() and helpers), gsd_input_val(), gsd_analysis_names().
  3. R/graph_test_shortcut_gsDesign.R — interface, normalization, validation, roxygen.
  4. R/print.gsd_graph_report.R — the gsDesign branch and gsd_spending_label().
  5. tests/testthat/test-gsd_engine.R, tests/testthat/test-graph_test_shortcut_gsDesign.R.
  6. vignettes/group-sequential-testing-gsDesign.Rmd.

Out of scope, noted for follow-up

  • repeated_p()/sequential_p() with spending_wt() in the mvtnorm path return 1 even at the final analysis (spending_wt() cannot calibrate its constant at alpha ≈ 1 and the error is swallowed); the gsDesign engine's "WT" path does not have this problem. Will be reported separately.
  • The original vignette's monitoring section says the analysis-1/2 monitoring boundaries differ from the planned ones; they are identical (spending preserved and the correlation between analyses 1 and 2 unchanged), as its own table shows. The new vignette states this correctly; the sentence in the original could be fixed separately.

🤖 Generated with Claude Code

xidongdxi and others added 2 commits September 18, 2026 15:44
Introduce gsd_engine_mvtnorm(), which bundles the only two group sequential computations the graphical procedure needs: repeated p-values and nominal p-value boundaries, backed by the existing repeated_p() and gs_boundaries(). gsd_test(), gsd_test_values_details(), gsd_test_values_look_back(), and gsd_boundary_table() now take an engine instead of spending_fn, so the graphical logic no longer depends on how those quantities are computed. graph_test_shortcut_gsd() builds the mvtnorm engine and passes it through; its interface and results are unchanged.

This is a behavior-neutral refactor that prepares for a gsDesign-backed graph_test_shortcut_gsDesign(). The full test suite passes unchanged (562 expectations, 0 failures), and new engine contract tests pin the interface a second engine must satisfy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
gs_boundaries() and repeated_p() integrate the joint distribution with randomized quasi-Monte Carlo (mvtnorm::GenzBretz), so two independent evaluations of the same boundary agree only to about 1e-5. The engine contract tests compared such pairs at expect_equal()'s default tolerance and failed on the third-analysis boundary.

Seed the RNG before each paired evaluation, allow a 1e-3 tolerance for the integration error, and additionally assert that the two hypotheses (O'Brien-Fleming vs Pocock spending) produce different results, which is the actual per-hypothesis dispatch guard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

badge

Code Coverage Summary

Filename                              Stmts    Miss  Cover    Missing
----------------------------------  -------  ------  -------  -----------------------------------------------
R/adjust_p.R                             44       1  97.73%   173
R/adjust_weights_parametric_util.R       55       0  100.00%
R/adjust_weights.R                       49       0  100.00%
R/as_graph.R                             37       2  94.59%   89, 105
R/edge_pairs.R                           11       1  90.91%   31
R/example_graphs.R                      181      15  91.71%   304, 408-422
R/graph_calculate_power.R               175       6  96.57%   378-383
R/graph_create.R                         75       0  100.00%
R/graph_generate_weights.R               44       0  100.00%
R/graph_rejection_orderings.R            35       0  100.00%
R/graph_test_closure.R                  181      10  94.48%   204-205, 300, 383-388, 401
R/graph_test_shortcut_gsd.R             362      13  96.41%   248, 610, 657, 716, 774, 793, 881-885, 887, 889
R/graph_test_shortcut_gsDesign.R        157      12  92.36%   206-207, 212-215, 315, 336-340
R/graph_test_shortcut.R                 106       0  100.00%
R/graph_update.R                         61       0  100.00%
R/gs_boundaries.R                        56       8  85.71%   60-61, 116-117, 120-121, 126-127
R/gs_corr.R                               4       0  100.00%
R/gsd_engine.R                          104       2  98.08%   157, 161
R/plot.initial_graph.R                   66       9  86.36%   135, 155, 161, 182-185, 202-205
R/plot.updated_graph.R                    4       0  100.00%
R/power_tests.R                          20       0  100.00%
R/print.graph_report.R                  177       1  99.44%   171
R/print.gsd_graph_report.R              188       5  97.34%   115-116, 118, 210, 322
R/print.initial_graph.R                  38       0  100.00%
R/print.power_report.R                  138       0  100.00%
R/print.updated_graph.R                  43       0  100.00%
R/repeated_p.R                           65       0  100.00%
R/sequential_p.R                         64       0  100.00%
R/spending_functions.R                  106       9  91.51%   145, 270-275, 433-434
R/test_power_input_val.R                122       0  100.00%
R/test_values.R                          80       0  100.00%
TOTAL                                  2848      94  96.70%

Diff against main

Filename                            Stmts    Miss  Cover
--------------------------------  -------  ------  -------
R/graph_test_shortcut_gsd.R           -13       0  -0.12%
R/graph_test_shortcut_gsDesign.R     +157     +12  +92.36%
R/gsd_engine.R                       +104      +2  +98.08%
R/print.gsd_graph_report.R            +32      +1  -0.10%
TOTAL                                +280     +15  -0.22%

Results for commit: f7f9f70

Minimum allowed coverage is 0%

♻️ This comment has been updated with latest results

@xidongdxi
xidongdxi marked this pull request as draft September 19, 2026 02:38
xidongdxi and others added 3 commits September 19, 2026 16:42
…graphical testing

Introduce a second group sequential engine, gsd_engine_gsDesign(), that obtains every group sequential quantity from the gsDesign package: one design per hypothesis via gsDesign(test.type = 1, n.I = info_frac, maxn.IPlan = 1, usTime), repeated p-values from sequentialPValue() with the earlier analyses' Z-statistics set to -20 (gsDesign's own no-lower-bound value), and nominal boundaries from the design's upper bound at the allocated level. The spending time is always passed explicitly because sequentialPValue()'s default rescales it by the last observed information. The classical boundary families ("OF", "Pocock", "WT") use a root-find on the design boundary, since sequentialPValue() does not support them; single-analysis hypotheses use the univariate closed form.

graph_test_shortcut_gsDesign() runs the same graphical procedure as graph_test_shortcut_gsd() on this engine, taking gsDesign's spending terms (sfu, sfupar, usTime) in place of spending_fn and returning the same gsd_graph_report structure. gsDesign remains a suggested package, checked at runtime. The print method labels gsDesign spending functions by their own names and shows usTime. gsd_input_val() now accepts a NULL spending_fn and the analysis-name logic is shared via gsd_analysis_names().

The two functions agree to about 1e-5 in p-values and boundaries (gsDesign integrates deterministically, mvtnorm by randomized quasi-Monte Carlo) with identical rejection decisions across the test scenarios, and the Wang-Tsiatis repeated p-values were cross-validated against rpact. Adds tests, roxygen, a NEWS entry, and the pkgdown reference entry.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Report lists of spending functions or parameters of the wrong length through the validation message instead of a names<- error, reject gsDesign::sfPoints() with a pointer to sfLinear() (sequentialPValue() evaluates the spending function on partial schedules, which sfPoints() refuses), store the supplied spending function in the design so that a user-defined function built on gsDesign::spendingFunction() is used consistently even without the sf self-reference, and keep list-valued parameters such as sfTruncated()'s out of the printed label. Documents that list parameters must be wrapped per hypothesis.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Mirrors group-sequential-testing.Rmd with graph_test_shortcut_gsDesign(): gsDesign's spending functions and their correspondence to the built-in ones, boundaries from gsDesign(), the sequentialPValue() mechanism with the earlier Z-statistics set to -20, both case studies, spending time via usTime, monitoring with usTime = planned and info_frac = actual (and the equivalent over-run representation), gsDesign's spending function library, and a user-defined spending function built on gsDesign::spendingFunction(). Knitting stops with a visible note if gsDesign is not installed.

Points to the new function from the original vignette's gsDesign subsection, adds the pkgdown article and a NEWS entry, and adds the new words to the word list.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@xidongdxi xidongdxi changed the title Delegate group sequential computations to an injectable engine Add graph_test_shortcut_gsDesign(): group sequential graphical testing backed by gsDesign Sep 19, 2026
@xidongdxi
xidongdxi marked this pull request as ready for review September 19, 2026 21:06
"Graphical Approaches in Group Sequential Design" for the original vignette and "Graphical Approaches in Group Sequential Design: Combining graphicalMCP and gsDesign" for the new one, so that the titles read as a pair and neither suggests that gsDesign handles the graphical part. File names and pkgdown URLs are unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant