HDAnalyzeR 1.1.0
A maintenance release focused on removing dependencies that no longer exist, getting continuous integration green again, and fixing the bugs that a rewritten test suite uncovered.
Breaking changes
-
multiROCandviphave been removed. Both were archived from CRAN (multiROCon 2026-05-04,vipon 2026-07-08), which made the package uninstallable from a clean library. Their functionality is now implemented in-package:- Multiclass ROC AUC is computed with
yardstick. The per-class one-vs-rest AUCs and themicroaverage are numerically identical to whatmultiROCreturned. Themacrovalue now uses the standard definition, the unweighted mean of the per-class AUCs, rather thanmultiROC’s curve-averaged variant, so it may differ slightly (by up to ~0.02 on small datasets). - Variable importance for the random forest and logistic regression engines is read directly from the fitted model and matches
vip::vi()exactly.
- Multiclass ROC AUC is computed with
-
Many packages moved from
ImportstoSuggests. Together with the removals below the package now installs with 28 hard dependencies instead of-
arrow,cluster,clusterProfiler,easyPubMed,embed,enrichplot,fpc,missForest,ppsr,readxl,WGCNAandwritexlare only needed by the specific functions that use them, and those functions now raise a clear error naming the package and the exact install command when it is missing. Install everything with:
# CRAN install.packages(c("arrow", "cluster", "easyPubMed", "embed", "fpc", "missForest", "ppsr", "readxl", "WGCNA", "writexl")) # Bioconductor BiocManager::install(c("clusterProfiler", "enrichplot", "org.Hs.eg.db", "ReactomePA")) -
purrr,tidyselectandumaphave been dropped entirely.umapwas never actually used:hd_umap()runs throughembed::step_umap(), which callsuwot. The check forumapinhd_umap()has been replaced with a check forembed.knitr,patchworkandviridismoved fromImportstoSuggests; they are only used by the vignettes.-
broom,forcats,readr,scales,stringrandtidytexthave been dropped. Each was pulled in for one or two small functions, which are now implemented inR/compat.Ror taken from a package that was already a hard dependency. Together with their own dependencies this removes 15 packages from a clean install (backports,bit,bit64,broom,clipr,crayon,forcats,hms,janeaustenr,progress,readr,SnowballC,tidytext,tokenizers,vroom), taking the install closure from 147 to 132 packages. There is no user-visible change in behaviour:-
broom::tidy()was only ever dispatching to methods registered byparsnipandrecipes, both already imported, so the calls now go through those. -
readr::read_csv()/read_tsv()inhd_import_data()are replaced withutils::read.table(), configured to keep readr’s semantics: feature names such asIL-6are not mangled, blank fields read asNA, and whole numbers read as double rather than integer. -
forcats::fct_reorder(),tidytext::reorder_within(),tidytext::scale_x_reordered(),tidytext::scale_y_reordered()andstringr::str_wrap()have base-R equivalents that produce identical output. -
scales::hue_pal()is replaced withgrDevices::hcl(). The palettes are identical up to 15 colours; beyond that one channel of one colour can differ by 1/255 becausescalesrounds viafarver.
-
rlangandwithrstay inImportseven though onlyrlang::sym()/syms()andwithr::local_seed()are used. Both remain in the install closure viadplyrandggplot2regardless, so replacing them would churn hundreds of call sites without removing a single package.glmnetandrangerstay inImports. They are the engineshd_model_rreg()andhd_model_rf()fit through, so the package cannot work without them even though it never calls them directly.extract_protein_list()(internal) no longer takes anUpSetRmembership matrix; it derives the combinations from the feature lists directly.
Bug fixes
Differential expression
-
hd_de_limma()no longer fails withlength of 'dimnames' [2] not equal to array extentwhencorrectcontains a categorical variable with more than two levels. Only the two group columns of the design matrix are renamed now. -
hd_de_ttest()returnedCI.L,CI.Randtas character strings, because each row was assembled withc()across mixed types. All statistics are numeric again. -
hd_plot_de_summary()labelled every feature"up"in bothproteins_df_upandproteins_df_down. The direction is now passed explicitly. -
hd_plot_de_summary()andhd_plot_model_summary()no longer error when summarising a single analysis.UpSetR::fromList()degenerates to a named vector for one set; the UpSet plot is now skipped with a message and returned asNULLwhen fewer than two groups have features.
Dimensionality reduction
-
hd_plot_dim(plot_loadings = ...)drew every loading arrow on the diagonal, because the tip used the same loading value for bothxendandyend. Arrows now run from the origin to the feature’s(x, y)loadings, and carry an arrowhead. -
hd_pca()returned loadings for every componentprcomp()produced rather than thecomponentsthat were requested, sopca_loadingsdisagreed withpca_resandpca_variance. - The point layer used an invalid
Colorlabel, which made ggplot2 emitIgnoring unknown labels.hd_plot_model_summary()had the same problem: its metrics bar plot labelledcolorwhile mappingfill, so the legend title was dropped.
Clustering
-
hd_cluster(normalize = TRUE)left the first feature completely unscaled and discarded the sample identifiers. It moved the sample ID into row names before callinghd_normalize(), which then treated the first feature as the ID column. Normalisation now happens before the row names are set. -
hd_cluster()returned the sample ID column as a factor; it is a character vector again. -
hd_assess_clusters()produced duplicated rows incluster_assessmentwhenever two clusters had the same size and stability, because the old and new cluster labels were matched on(n, Mean_ji). They are matched on the original cluster label now.
Enrichment
-
hd_gsea(ranked_by = "both")computed the combined logFC / significance score and then ranked byadj.P.Valanyway. It now ranks by the score it computed. -
hd_gsea()renamed the ranked gene vector to ENTREZ identifiers by position. Because symbol mapping drops unmappable genes, every score after the first failure was attached to the wrong gene. Scores are now matched by symbol. -
hd_gsea()is reproducible. GSEA p-values come from a permutation test, so every call returned a slightly different set of enriched terms. A newseedargument (123 by default,NULLto opt out) fixes the random number generator for the duration of the call. -
hd_ora()andhd_gsea()no longer stop when nothing is enriched. They warn and return the empty enrichment object instead. A null result is a legitimate outcome, and aborting on it broke the vignettes and the examples in CI whenever the permutation test or a Bioconductor annotation update happened to push everything above the threshold.hd_plot_ora()andhd_plot_gsea()warn and return their input unchanged in that case, rather than failing insideclusterProfiler. -
hd_ora()tested for missing adjusted p-values withis.na(any(p.adjust)), which coerces the numbers to logicals and never detects anything. An all-NAcolumn now reports that nothing was enriched instead of failing withmissing value where TRUE/FALSE needed. - The returned
hd_enrichmentobject carries thepval_limit was built with, so the plotting functions know which terms were considered significant.
Co-expression networks
-
hd_wgcna()passedNAstraight toWGCNA::blockwiseModules()whenpickSoftThreshold()found no power meeting the scale-free topology cut-off. It now falls back to the WGCNA default power of 6 and says so. -
hd_plot_wgcna()errored whenclinical_varscontained only continuous variables. -
hd_plot_wgcna()bound the module eigengenes and the metadata side by side, silently pairing modules with the wrong samples whenever the metadata rows were in a different order. They are joined on the sample ID now.
Models
-
hd_model_rf()failed withThe 'range' lower bound (2) must not exceed upper bound (1)on datasets with fewer than about nine predictors, becausemtrywas tuned overc(floor(sqrt(p)), floor(p / 3)). The range is clamped now.
Utilities and reporting
-
hd_import_data()returned one of the function’s own local variables instead of the stored object when reading.rdafiles, because it usedls()[1]. It uses the names returned byload()now. -
hd_save_data()corrupted values containing a double-quote when writing TSV.utils::write.table()defaults to escaping an embedded quote as\", which no reader parses back, sohe said "hi"reimported ashe said \hi\. Quotes are doubled now, the wayutils::write.csv()already did for CSV. -
hd_qc_summary(),hd_impute_median(),hd_impute_knn()andhd_impute_missForest()printed tables as deparsed R code (c("f2", "f1")c(1, 1)). Tables are rendered the way they print at the console, truncated to the first ten rows. -
hd_filter()andhd_save_path()produced messages with no spaces between the words and the values (Rows remaining:3,Directoryoutalready exists.). -
hd_normalize()leftscaled:centerandscaled:scaleattributes on every column of the returned tibble.
Performance
hd_qc_summary() ran out of memory on large datasets (10,000 features, 5,000 samples). The results are unchanged on every input; only the way they are computed differs.
-
The correlation heatmap is capped.
hd_plot_cor_heatmap()andhd_qc_summary()gainedmax_heatmap_features(default 1000). Above the limit the correlation matrix and the reported pairs are still returned, butcor_heatmapisNULLand a warning explains why. Clustering 10,000 features needs a 400 MB distance matrix and hours ofhclust, for a plot with 100 million unreadable cells. -
The high-correlation pairs are read straight off the matrix. The pairs used to be found by reshaping the whole correlation matrix to long format first, which is one row per entry: 100 million rows before any filtering, several GB. The new
cor_pairs_above()scans blocks of columns and keeps only what exceeds the threshold, in the same order as before. -
calc_na_percentage_row()no longer usesrowwise(), which evaluated once per sample. Counting withrowSums()over blocks of columns is roughly 2000 times faster on a 1,000 × 3,000 dataset (21 s to 0.01 s) and uses a bounded amount of memory. -
calc_na_percentage_col()counts a column at a time instead of building a one-row, 10,000-column summary and pivoting it. -
hd_correlate()substitutesuse = "everything"for"pairwise.complete.obs"when the input has no missing values. The two are equivalent in that case, and the pairwise code path is about three times slower because it compares every pair of columns separately. -
check_numeric_columns()no longer coerces already-numeric columns withas.numeric(), which allocated a copy of every column of the dataset. -
hd_qc_summary()now says up front when a correlation is going to take a while.
Continuous integration
- Fixed
.Rbuildignore: the patterns forinst/extdata,inst/cheatsheetandinst/hdanalyzer_appused\e,\cand\hinstead of/, so they never matched. The source tarball was 49 MB; it is now 604 KB. -
R-CMD-checkalso runs ondev/**andka/**branches and onworkflow_dispatch, not only onmain. Pushes to working branches were never building. - Added
bioc-versionto the R setup so the Bioconductor dependencies resolve reliably, andoldrel-1to the test matrix. - The R-devel job is now
continue-on-error. Bioconductor routinely lags R-devel for weeks after a release, which used to fail the whole matrix for reasons unrelated to this package. - Added a job timeout and
concurrencycancellation so superseded runs stop instead of queueing. - Fixed a malformed
Suggestsfield inDESCRIPTION(a trailing comma) and a redirecting URL. - The
hd_literature_search()example queried PubMed live duringR CMD check, so a throttled request could stall the examples step on every platform. It is wrapped in\donttest{}now. - The
hd_gsea()examples and thepost_analysisvignette failed the examples and pkgdown steps withNo significant terms found. GSEA is a permutation test, so whether anything cleared the threshold varied between machines and between runs.hd_gsea()is seeded and no longer treats an empty result as an error (see Enrichment above). - Every vignette declared a
\VignetteIndexEntry{}that did not match its YAML title, so each build printed a warning about it. The index entries now carry the real titles. -
Examples that need a
Suggestspackage are now conditional. Twelve help topics ran examples that call intoclusterProfiler,org.Hs.eg.db,enrichplot,WGCNA,ppsr,missForest,cluster,fpc,embedoreasyPubMedunconditionally. Suggested packages are best-effort on CI — the Windows runner had no usableclusterProfiler, andR CMD checkthen failed withThe 'clusterProfiler' package is required ... but is not installedwhile macOS and Ubuntu passed. Each of those topics now carries an@examplesIf requireNamespace(...)guard, so the examples run where the package is available and are skipped where it is not. -
The pkgdown workflow survives, and reports, a crash in the article subprocess. pkgdown renders each article with
callr::r_safe()and polls the subprocess every 200 ms until it exits. The site build died twice while polling the render ofclassification.Rmd, the longest article, inside processx’s own internals —chain_call()on one run,assert_that()on the next. pkgdown cannot format that condition because it carries no$stderr, so it reportedsubscript out of boundsfromwrap_rmarkdown_error()and the real cause never reached the log. This is not in the package: every article renders cleanly in-process on the same runner, which has around 14 GB of memory and 79 GB of disk free. The workflow now:- installs the newest
callrandprocessxrather than whichever version is cached, since that is where the crash happens; - builds with
quiet = FALSE, so the subprocess output is streamed into the log and a genuine article error stays readable; - retries once with
clean = FALSE, lazy = TRUEon failure, which resumes from the articles that already rendered instead of starting over; - reports the runner’s memory and disk, and re-renders every article in-process, both only when the build fails.
- installs the newest
- Added build and check artefacts (
*.Rcheck/,*.tar.gz,Rplots.pdf) to.gitignore.
Testing
- The test suite has been rewritten. It now checks behaviour and numeric results rather than only object shapes: differential expression is verified against
stats::t.test, imputation and normalisation against hand-computed values, multiclass AUC against a one-vs-rest computation, and variable importance against the model coefficients. - Added end-to-end tests covering the full workflows on the shipped example data: QC to imputation to normalisation to PCA to differential expression to classification, plus the multi-class and multi-model summary paths.
- Tests for optional packages are skipped rather than failed when the package is not installed.
- The PubMed tests are opt-in. NCBI throttles unauthenticated clients and a throttled request can stall for minutes, which is exactly the kind of hang that makes CI look broken. Run them with
HDANALYZER_TEST_PUBMED=true. - The end-to-end tests run on a five-disease, forty-assay subset of the example data. They cover the same code paths as the full dataset but the whole suite finishes in about a minute, so it is cheap enough to run on every push.
The suite is 15 files, 229 test_that() blocks and 608 assertions, passing with no failures.
R CMD check --as-cran now completes with 0 errors and 0 warnings. Two NOTEs remain and neither is a defect: “New submission” (the package is not on CRAN) and a handful of examples running longer than five seconds, which is inherent to fitting real models on the example data.
Internal
- Replaced superseded verbs (
summarise_all(),gather(),spread(),top_n(),sample_n(),rename_all()) with their current tidyverse equivalents. - Fixed
tidyselect“external vector in selections” deprecation warnings inhd_split_data()and the model tuning helpers. - Replaced the deprecated
sizeaesthetic withlinewidthinhd_show_palettes(). - Removed 16 assigned-but-unused
check_numeric_columns()results and other dead code. - Extracted
multiclass_auc(),model_importance(),rank_features(),rename_ranking_to_entrezid(),build_upset_plot(),message_table()andcheck_installed()as documented internal helpers, which makes the previously untestable logic directly testable.
