Reads tabular data from xlsx files with a specialized C parser. Worksheet XML is scanned in a single pass and decoded directly into R vectors, with no intermediate document model. Bundles the 'miniz' and 'libdeflate' decompressors to read the underlying archive.
Read tabular data from xlsx files with a native C parser.
rcxl decodes each worksheet directly into R vectors as it is scanned, with no
intermediate document model. This keeps memory use close to the size of the
data itself. The miniz and libdeflate decompressors are bundled, so the
package has no R dependencies beyond base R.
From CRAN:
install.packages("rcxl")
The development version from GitHub:
# install.packages("remotes")
remotes::install_github("vlshields/rcxl")
read_xlsx() returns a plain data.frame, taking the first row as the header:
library(rcxl)
path <- system.file("extdata", "flights.xlsx", package = "rcxl")
df <- read_xlsx(path)
dim(df)
#> [1] 100 10
Column types are guessed from the cells in the read area; dates come back as
POSIXct in UTC.
xlsx_sheets() lists sheets, and sheet takes a name or an index:
multi <- system.file("extdata", "multisheet.xlsx", package = "rcxl")
xlsx_sheets(multi)
#> [1] "Alpha" "Beta" "Gamma"
read_xlsx(multi, sheet = "Beta")
read_xlsx(multi, sheet = 2)
read_xlsx_all() reads every sheet from a single workbook parse and returns a
named list of data.frames:
sheets <- read_xlsx_all(multi)
names(sheets)
#> [1] "Alpha" "Beta" "Gamma"
Reads can be windowed with an A1 range (column-only and row-only forms work
too), or with skip / n_max:
read_xlsx(path, range = "B3:D10")
read_xlsx(path, range = "B:C")
read_xlsx(path, skip = 5, n_max = 10)
col_names accepts TRUE, FALSE (auto names, header row becomes data), or a
character vector sized to the read area. col_types takes "text",
"numeric", "date", "logical", "guess", "skip", or "list", recycled
if length 1:
read_xlsx(path, col_names = FALSE)
read_xlsx(path, col_types = "text") # everything as character
read_xlsx(path, col_types = c("skip", rep("guess", 9))) # drop the first column
Other arguments: na blanks matching cells, trim_ws trims surrounding
whitespace, and name_repair is one of "unique" (default), "minimal", or
"check_unique". See ?read_xlsx for details.
Set the RCXL_THREADS environment variable to override the worker count.
Sheets under 4MB are read serially regardless.
MIT. Bundled miniz and libdeflate retain their own copyright and license
terms; see inst/COPYRIGHTS.