Fits Type-1 Fuzzy Function models for binary classification, numeric regression, and time-series forecasting with user-supplied temporal predictors. The package combines fuzzy C-means memberships, nonlinear membership transformations, cluster-specific linear or support vector machine models, and membership-weighted predictions. It also provides model evaluation, validation, K-fold and stratified K-fold tuning, and repeated nested cross-validation with task-appropriate metrics. The regression workflow can be used for forecasting when temporal dependence is represented by lagged or seasonal predictors and assessment partitions preserve chronological order.
T1FF fits Type-1 Fuzzy Function models for binary classification, numeric
regression, and time-series forecasting with user-supplied temporal predictors.
It uses fuzzy C-means memberships, cluster-specific statistical models, and
membership-weighted aggregation. Forecasting requires time dependence to be
encoded with lagged or seasonal predictors and training, validation, and test
sets to preserve chronological order.
# Install the local source package
install.packages("T1FF_0.1.0.tar.gz", repos = NULL, type = "source")
# After a CRAN release
# install.packages("T1FF")
library(T1FF)
data(iris)
d <- droplevels(subset(iris, Species != "setosa"))
fit <- T1FF(
da = d,
target_col = "Species",
c = 2,
m = 2,
task = "classification",
positive_class = "virginica",
seed = 123
)
probability <- predict(fit, d, type = "prob")
class_hat <- predict(fit, d, type = "class")
cluster_probability <- predict(fit, d, type = "cluster_prob")
membership <- predict(fit, d, type = "membership")
For classification, logistic_method = "auto" is the default. It detects
separation or extreme local GLM fits and refits affected local models with
ridge-penalized logistic regression. ridge_lambda controls the penalty and
probability_clip (default 1e-6) prevents exact zero/one probabilities.
fit <- T1FF(
d, "Species", c = 2, task = "classification",
positive_class = "virginica", logistic_method = "auto",
ridge_lambda = 0.01, probability_clip = 1e-6, seed = 1
)
summary(fit)
Use probabilistic support-vector machines within the fuzzy clusters with
local_model = "svm":
fit_svm <- T1FF(
d, "Species", c = 2, m = 2,
task = "classification", local_model = "svm",
svm_kernel = "rbfdot", svm_C = 1,
positive_class = "virginica", seed = 1
)
predict(fit_svm, d[1:5, ], type = "prob")
Evaluate a fitted or tuned model with task-appropriate metrics:
evaluation <- evaluate.T1FF(fit, d, truth = "Species")
summary(evaluation)
evaluation$metrics
evaluation$confusion_matrix
Estimate generalization performance with repeated nested cross-validation. The inner folds tune T1FF while the outer folds compare it with a standard logistic or linear regression baseline:
benchmark <- benchmark.T1FF(
Species ~ ., d,
c_values = 2:4, m_values = c(1.5, 2, 2.5),
outer_folds = 5, inner_folds = 5, repeats = 2,
positive_class = "virginica", seed = 1
)
summary(benchmark)
benchmark$fold_results
benchmark$selected_parameters
benchmark$timing_summary
The equivalent formula interface also supports factor and character predictors through a training-derived design matrix:
fit <- T1FF(
Species ~ .,
data = d,
c = 2,
task = "classification",
positive_class = "virginica",
na_action = "fail",
seed = 123
)
Set na_action = "omit" to remove incomplete training rows. Predictions for
incomplete rows are then returned as NA, preserving the row count and order
of newdata.
fit <- T1FF(
da = mtcars,
target_col = "mpg",
c = 2,
m = 2,
task = "regression",
seed = 123
)
prediction <- predict(fit, mtcars, type = "response")
cluster_prediction <- predict(fit, mtcars, type = "cluster_response")
membership <- predict(fit, mtcars, type = "membership")
Regression tuning and evaluation support rmse, mse, mae, mape,
smape, and r2. MAPE and SMAPE are returned as percentages; zero actual
values are excluded from MAPE, while a jointly zero actual and prediction has
zero contribution to SMAPE.
tuned_regression <- tune.T1FF(
mtcars, "mpg", task = "regression",
c_values = 2:3, m_values = c(1.5, 2),
metric = "smape", resampling = "kfold", folds = 5,
seed = 123
)
evaluate.T1FF(tuned_regression, mtcars, truth = "mpg")
T1FF does not automatically infer a time index or create lagged predictors.
For forecasting, first express temporal dependence through variables such as
response lags, seasonal terms, or a trend. Always train on earlier observations
and evaluate on later observations. The built-in tuning folds are random, so
time-series hyperparameter selection should use an external chronological
validation or rolling-origin procedure.
This example performs rolling one-step-ahead evaluation on AirPassengers;
the test-row lags contain only values observed by each forecast origin.
data(AirPassengers)
y <- as.numeric(AirPassengers)
period <- as.numeric(time(AirPassengers))
month <- as.numeric(cycle(AirPassengers))
n <- length(y)
forecast_data <- data.frame(
period = period[13:n],
y = y[13:n],
lag1 = y[12:(n - 1)],
lag12 = y[1:(n - 12)],
trend = seq_len(n - 12),
season_sin = sin(2 * pi * month[13:n] / 12),
season_cos = cos(2 * pi * month[13:n] / 12)
)
split <- floor(0.80 * nrow(forecast_data))
train_ts <- forecast_data[seq_len(split), ]
test_ts <- forecast_data[(split + 1):nrow(forecast_data), ]
forecast_fit <- T1FF(
y ~ lag1 + lag12 + trend + season_sin + season_cos,
data = train_ts, c = 2, m = 2, task = "regression", seed = 123
)
test_ts$forecast <- predict(forecast_fit, test_ts, type = "response")
evaluate.T1FF(forecast_fit, test_ts, truth = "y")
tuned <- tune.T1FF(
da = d,
target_col = "Species",
task = "classification",
c_values = 2:3,
m_values = c(1.5, 2),
metric = "logloss",
resampling = "stratified_kfold",
folds = 5,
positive_class = "virginica",
seed = 123
)
summary(tuned)
predict(tuned, d, type = "prob")
Classification tuning supports logloss, brier, roc_auc, pr_auc,
accuracy, balanced_accuracy, f1, sensitivity, and specificity.
Tune the classification threshold together with c and m for a
threshold-dependent metric:
tuned_threshold <- tune.T1FF(
Species ~ ., d,
c_values = 2:4,
m_values = c(1.5, 2, 2.5),
threshold_values = seq(0.3, 0.7, by = 0.05),
metric = "balanced_accuracy",
folds = 5,
positive_class = "virginica",
seed = 123
)
tuned_threshold$best_threshold
predict(tuned_threshold, d, type = "class")
The package is under active development. The first release supports numeric, factor, and character predictors; binary outcomes for classification; numeric outcomes for regression; and explicit missing-value handling.