Provides tools to quantify how strongly pairs of words attract or repel each other in a text corpus, based on co-occurrence patterns. For each word pair, the phi coefficient (a correlation measure for binary variables) is computed from a document-term matrix and tested for significance, then classified as showing attraction (co-occurring more than chance would predict), repulsion (co-occurring less than chance would predict), or no significant relationship. A full pipeline is provided from raw text to a labeled network visualization. Unlike general-purpose pairwise correlation tools, 'wordorientation' is built specifically for text: it handles tokenization and stopword removal, applies significance-based classification rather than reporting a raw correlation coefficient alone, and produces a ready-to-plot attraction/ repulsion network.
Detect attraction and repulsion between words in text.
For every pair of words that co-occur often enough to analyze,
wordorientation computes the phi coefficient (a correlation measure for
binary co-occurrence data), tests it for significance, and classifies the
pair as:
# once on CRAN
install.packages("wordorientation")
# development version
# devtools::install_github("yourusername/wordorientation")
library(wordorientation)
result <- analyze_word_orientation(
example_social_posts(),
text_col = "text", doc_col = "id",
min_count = 2
)
head(result$scored)
plot_orientation_network(result$scored)
Or step by step:
tokens <- tokenize_posts(my_data, text_col = "text")
cooc <- cooccurrence_counts(tokens, min_count = 5)
scored <- word_orientation(cooc, alpha = 0.05)
widyr::pairwise_cor() computes the same underlying phi-style
correlation but is a general tidy-correlation tool, not text-specific: it
has no built-in tokenization, no significance-based classification, and
no network plotting.collostructions measures the attraction/repulsion of words to
grammatical constructions, not to each other.MadanTextNetwork provides a co-occurrence network Shiny app but is
built specifically for Persian text and does not classify pairs by
statistical significance.wordorientation combines tokenization, co-occurrence counting,
significance-tested classification, and network visualization into a single
general-language pipeline.
By default, word_orientation() applies a Benjamini-Hochberg correction
across all tested word pairs (p_adjust_method = "BH"). On small corpora
with few documents, this can mean no pair survives correction even when raw
phi values look large — this is intentional, not a bug: it guards against
over-interpreting spurious associations from sparse data.