diff --git a/DESCRIPTION b/DESCRIPTION index c571bab747..97fb693cc4 100644 --- a/DESCRIPTION +++ b/DESCRIPTION @@ -110,6 +110,4 @@ Authors@R: c( person("Manmita", "Das", role="ctb"), person("Tarun", "Thammisetty", role="ctb"), person("Marco", "Colombo", role="ctb", comment = c(ORCID = "0000-0001-6672-0623")), - person("Tim", "Taylor", role="ctb", comment = c(ORCID = "0000-0002-8587-7113")), - NULL - ) + person("Tim", "Taylor", role="ctb", comment = c(ORCID = "0000-0002-8587-7113"))) diff --git a/NEWS.md b/NEWS.md index 61a526de41..ee2f96fd97 100644 --- a/NEWS.md +++ b/NEWS.md @@ -102,7 +102,7 @@ 5. `melt()` and `dcast()` no longer provide nudges when receiving incompatible inputs (e.g. data.frames). As of now, we only define methods for `data.table` inputs. -6. Enhanced tests for OpenMP support, detecting incompatibilities such as R-bundled runtime _vs._ newer Xcode and testing for a manually installed runtime from , [#6622](https://github.com/Rdatatable/data.table/issues/6622). Thanks to @dvg-p4 for initial report and testing, @twitched for the pointers, @tdhock and @aitap for the fix. +6. Enhanced tests for OpenMP support, detecting incompatibilities such as R-bundled runtime _vs._ newer Xcode and testing for a manually installed runtime from , [#6622](https://github.com/Rdatatable/data.table/issues/6622). Thanks to @dvg-p4 for initial report and testing, @twitched for the pointers, @tdhock and @aitap for the fix. 7. Verbose outputs from `frolladaptivefun()` and `frollfun()` are now clearer and more user friendly [#7021](https://github.com/Rdatatable/data.table/issues/7021). Thanks to @Omartech312, @aidengseay, @kkarissa, and @heb229 for the implementation, to @ben-schwen for the review, and to @jangorecki for the extensive guidance and review. @@ -125,7 +125,7 @@ 2. `rbindlist()` (and therefore the `rbind()` method for `data.table`s) no longer raises an error upon encountering more than approximately 50000 columns in a list entry, [#7793](https://github.com/Rdatatable/data.table/issues/7793). The bug was introduced in `data.table` version 1.18.2.1. Thanks to @rickhelmus for the report and @aitap for the fix. -## NOTES +### NOTES 1. Handled OpenMP deprecation of `master` construct, [#7882](https://github.com/Rdatatable/data.table/pull/7882). Thanks @TimTaylor for the PR. diff --git a/README.md b/README.md index 0f6dddf5d8..3dbbbfdd09 100644 --- a/README.md +++ b/README.md @@ -6,7 +6,7 @@ [![R-CMD-check](https://github.com/Rdatatable/data.table/actions/workflows/R-CMD-check.yaml/badge.svg?branch=master)](https://github.com/Rdatatable/data.table/actions) [![Codecov test coverage](https://codecov.io/github/Rdatatable/data.table/coverage.svg?branch=master)](https://app.codecov.io/github/Rdatatable/data.table?branch=master) [![GitLab CI build status](https://gitlab.com/Rdatatable/data.table/badges/master/pipeline.svg)](https://rdatatable.gitlab.io/data.table/web/checks/check_results_data.table.html) -[![downloads](https://cranlogs.r-pkg.org/badges/data.table)](https://www.rdocumentation.org/trends) +[![downloads](https://cranlogs.r-pkg.org/badges/data.table)](https://github.com/r-hub/cranlogs.app) [![CRAN usage](https://jangorecki.gitlab.io/rdeps/data.table/CRAN_usage.svg?sanitize=true)](https://gitlab.com/jangorecki/rdeps) [![BioC usage](https://jangorecki.gitlab.io/rdeps/data.table/BioC_usage.svg?sanitize=true)](https://gitlab.com/jangorecki/rdeps) [![indirect usage](https://jangorecki.gitlab.io/rdeps/data.table/indirect_usage.svg?sanitize=true)](https://gitlab.com/jangorecki/rdeps) diff --git a/man/assign.Rd b/man/assign.Rd index 1eb9daa674..cba7d6f87b 100644 --- a/man/assign.Rd +++ b/man/assign.Rd @@ -32,6 +32,11 @@ # DT[i, names(.SD) := lapply(.SD, fx), by = ..., .SDcols = ...] set(x, i = NULL, j, value) + +# Exported as: +`:=`(...) +let(...) +# Please don't call from outside j-expressions on data.tables. } \arguments{ \item{LHS}{ A character vector of column names (or numeric positions) or a variable that evaluates as such. If the column doesn't exist, it is added, \emph{by reference}. } @@ -42,6 +47,7 @@ set(x, i = NULL, j, value) In \code{set}, only integer type is allowed in \code{i} indicating which rows \code{value} should be assigned to. \code{NULL} represents all rows more efficiently than creating a vector such as \code{1:nrow(x)}. } \item{j}{ Column name(s) (character) or number(s) (integer) to be assigned \code{value} when column(s) already exist, and only column name(s) if they are to be created. } \item{value}{ A list or vector of replacement values to be assigned by reference to \code{x[i, j]}. } +\item{...}{Ignored by the exported functions \code{`:=`} and \code{let} when called from outside a \code{j}-expression.} } \details{ \code{:=} is defined for use in \code{j} only. It \emph{adds} or \emph{updates} or \emph{removes} column(s) by reference. It makes no copies of any part of memory at all. Please read \href{../doc/datatable-reference-semantics.html}{\code{vignette("datatable-reference-semantics")}} and follow with examples. Some typical usages are: diff --git a/man/dcast.data.table.Rd b/man/dcast.data.table.Rd index 6187dfce79..87dfb27e09 100644 --- a/man/dcast.data.table.Rd +++ b/man/dcast.data.table.Rd @@ -7,11 +7,13 @@ } \usage{ +dcast(data, formula, fun.aggregate = NULL, ..., margins = NULL, + subset = NULL, fill = NULL, value.var = guess(data)) \method{dcast}{data.table}(data, formula, fun.aggregate = NULL, sep = "_", \dots, margins = NULL, subset = NULL, fill = NULL, drop = TRUE, value.var = guess(data), verbose = getOption("datatable.verbose"), - value.var.in.dots = FALSE, value.var.in.LHSdots = value.var.in.dots, + value.var.in.dots = FALSE, value.var.in.LHSdots = value.var.in.dots, value.var.in.RHSdots = value.var.in.dots) } \arguments{ diff --git a/man/melt.data.table.Rd b/man/melt.data.table.Rd index 40506eb6ab..b626f50287 100644 --- a/man/melt.data.table.Rd +++ b/man/melt.data.table.Rd @@ -9,6 +9,7 @@ efficiency. Since \code{v1.9.6}, \code{melt.data.table} allows melting into multiple columns simultaneously. } \usage{ +melt(data, ..., na.rm = FALSE, value.name = "value") ## fast melt a data.table \method{melt}{data.table}(data, id.vars, measure.vars, variable.name = "variable", value.name = "value", diff --git a/man/nafill.Rd b/man/nafill.Rd index ef61554db8..b7450529f8 100644 --- a/man/nafill.Rd +++ b/man/nafill.Rd @@ -11,7 +11,10 @@ } \usage{ nafill(x, type=c("const", "locf", "nocb"), fill=NA, nan=NA, limit=Inf) -setnafill(x, type=c("const", "locf", "nocb"), fill=NA, nan=NA, cols=seq_along(x), limit=Inf) +setnafill( + x, type=c("const", "locf", "nocb"), fill=NA, nan=NA, cols=seq_along(x), + limit=Inf +) } \arguments{ \item{x}{ Vector, list, data.frame or data.table of logical, numeric or character columns. } diff --git a/vignettes/datatable-faq.Rmd b/vignettes/datatable-faq.Rmd index 406dbdaf06..0aa603561c 100644 --- a/vignettes/datatable-faq.Rmd +++ b/vignettes/datatable-faq.Rmd @@ -525,7 +525,7 @@ copied in bulk (`memcpy` in C) rather than looping in C. ## What are primary and secondary indexes in data.table? -Manual: [`?setkey`](https://www.rdocumentation.org/packages/data.table/functions/setkey) +Manual: [`?setkey`](https://r-datatable.com/reference/setkey.html) S.O.: [What is the purpose of setting a key in data.table?](https://stackoverflow.com/questions/20039335/what-is-the-purpose-of-setting-a-key-in-data-table/20057411#20057411) `setkey(DT, col1, col2)` orders the rows by column `col1` then within each group of `col1` it orders by `col2`. This is a _primary index_. The row order is changed _by reference_ in RAM. Subsequent joins and groups on those key columns then take advantage of the sort order for efficiency. (Imagine how difficult looking for a phone number in a printed telephone directory would be if it wasn't sorted by surname then forename. That's literally all `setkey` does. It sorts the rows by the columns you specify.) The index doesn't use any RAM. It simply changes the row order in RAM and marks the key columns. Analogous to a _clustered index_ in SQL.