Read xlsx Files with a Native Parser

Reads tabular data from xlsx files with a specialized C parser. Worksheet XML is scanned in a single pass and decoded directly into R vectors, with no intermediate document model. Bundles the 'miniz' and 'libdeflate' decompressors to read the underlying archive.


rcxl

CRAN status R-CMD-check

Read tabular data from xlsx files with a native C parser.

rcxl decodes each worksheet directly into R vectors as it is scanned, with no intermediate document model. This keeps memory use close to the size of the data itself. The miniz and libdeflate decompressors are bundled, so the package has no R dependencies beyond base R.

Installation

From CRAN:

install.packages("rcxl")

The development version from GitHub:

# install.packages("remotes")
remotes::install_github("vlshields/rcxl")

Usage

read_xlsx() returns a plain data.frame, taking the first row as the header:

library(rcxl)

path <- system.file("extdata", "flights.xlsx", package = "rcxl")
df <- read_xlsx(path)
dim(df)
#> [1] 100  10

Column types are guessed from the cells in the read area; dates come back as POSIXct in UTC.

xlsx_sheets() lists sheets, and sheet takes a name or an index:

multi <- system.file("extdata", "multisheet.xlsx", package = "rcxl")
xlsx_sheets(multi)
#> [1] "Alpha" "Beta"  "Gamma"

read_xlsx(multi, sheet = "Beta")
read_xlsx(multi, sheet = 2)

read_xlsx_all() reads every sheet from a single workbook parse and returns a named list of data.frames:

sheets <- read_xlsx_all(multi)
names(sheets)
#> [1] "Alpha" "Beta"  "Gamma"

Reads can be windowed with an A1 range (column-only and row-only forms work too), or with skip / n_max:

read_xlsx(path, range = "B3:D10")
read_xlsx(path, range = "B:C")
read_xlsx(path, skip = 5, n_max = 10)

col_names accepts TRUE, FALSE (auto names, header row becomes data), or a character vector sized to the read area. col_types takes "text", "numeric", "date", "logical", "guess", "skip", or "list", recycled if length 1:

read_xlsx(path, col_names = FALSE)
read_xlsx(path, col_types = "text")                    # everything as character
read_xlsx(path, col_types = c("skip", rep("guess", 9)))  # drop the first column

Other arguments: na blanks matching cells, trim_ws trims surrounding whitespace, and name_repair is one of "unique" (default), "minimal", or "check_unique". See ?read_xlsx for details.

Set the RCXL_THREADS environment variable to override the worker count. Sheets under 4MB are read serially regardless.

License

MIT. Bundled miniz and libdeflate retain their own copyright and license terms; see inst/COPYRIGHTS.

Reference manual

It appears you don't have a PDF plugin for this browser. You can click here to download the reference manual.

install.packages("rcxl")

0.1.1 by Vincent Shields, 3 days ago


https://github.com/vlshields/rcxl


Report a bug at https://github.com/vlshields/rcxl/issues


Browse source code at https://github.com/cran/rcxl


Authors: Vincent Shields [aut, cre, cph] , Rich Geldreich [ctb, cph] (author of the bundled miniz) , Tenacious Software LLC [cph] (copyright holder of the bundled miniz) , RAD Game Tools and Valve Software [cph] (copyright holder of the bundled miniz) , Martin Raiber [ctb, cph] (contributor to and copyright holder of the miniz ZIP reader) , Alex Evans [ctb] (author of the PNG writer in the bundled miniz) , Alistair Moffat [ctb] (author of the minimum-redundancy code in the bundled miniz) , Jyrki Katajainen [ctb] (author of the minimum-redundancy code in the bundled miniz) , Eric Biggers [ctb, cph] (author of the bundled libdeflate) , Google LLC [cph] (copyright holder of the bundled libdeflate)


Documentation:   PDF Manual  


MIT + file LICENSE license


Suggests readxl, cellranger


See at CRAN