Fine-tune small (100M to 3B parameter) causal language models with LoRA (Low-Rank Adaptation) from R. Datasets are mapped to chat-format prompts and responses, training runs in a background 'Python' process built on Hugging Face 'transformers' and 'peft', and a 'shiny' app offers drag-and-drop dataset upload and column mapping. 'Python' dependencies are declared through 'reticulate' and resolved automatically on first use.
Fine-tune small language models with LoRA from R, by code or by drag and drop.
dragonfarm takes a table of prompts and replies, teaches a 100M to 3B parameter
model to answer in your format and your domain, and gives you back an adapter
or a merged model that loads with plain Hugging Face transformers. Training
runs in a background Python process that the package sets up for you.
# install.packages("pak")
pak::pak("tejas4patel/dragon-farm")
Then check the machine. The first call builds a Python environment with torch and transformers, which downloads 2 to 3 GB and takes a few minutes.
library(dragonfarm)
dragon_check()
reticulate uses a Python 3.10 to 3.13 it
finds on the machine, or downloads one. torch, transformers, and peft are
installed automatically on first use.nvidia-smi.Run dragon_check() after installing. It reports the device it will train
on, and if that is the CPU on a machine with an NVIDIA GPU it says why (no
driver, a driver too old for the installed torch, or a CPU-only torch build)
and prints the one-line fix.
Where runs are stored. By default a run's files (adapter, checkpoints,
data, logs) go under a dragonfarm_runs folder inside a session temp
directory, so a fresh R session never writes to your working directory or
home filespace on its own; that folder disappears when the session ends.
For runs you want to keep, set a real location once per project, before
training anything:
options(dragonfarm.runs_dir = "~/dragonfarm_runs") # or any path you like
or set the DRAGONFARM_RUNS_DIR environment variable, or pass runs_dir =
to dragon_train(), dragon_bundle(), dragon_app(), and the rest
directly. See ?dragon_runs_dir.
library(dragonfarm)
run <- dragon_dataset(dragon_example_data()) |>
dragon_map(prompt = "{subject}\n\n{body}", response = "reply") |>
dragon_train("Qwen/Qwen2.5-0.5B-Instruct", wait = TRUE)
dragon_generate(run, "My thermostat keeps dropping off Wi-Fi.")
dragon_merge(run, "models/support-0.5b")
dragon_train() returns immediately by default. Runs live on disk, so you can
close R and come back:
run <- dragon_run("dragonfarm_runs/20260913-143201-qwen2.5-0.5b-instruct")
dragon_status(run)
dragon_progress(run) # one row per logged step
dragon_wait(run) # progress bar until it finishes
dragon_cancel(run) # stops after the current step and saves a checkpoint
dragon_resume(run) # picks up from that checkpoint
dragon_app()
Six panels, left to right: drop a file, drag its columns into Prompt and Response slots, pick a model, set a few numbers, watch the loss curve, and compare the tuned model against the base model. Every run started in the app is a normal run directory, and the Monitor panel shows the R code that reproduces it.
Fine-tuning teaches the model what a good reply looks like. The next stage teaches it which of two replies is better, from a table with a prompt, a chosen reply, and a rejected one. It runs on top of a fine-tuned run:
sft <- dragon_dataset("tickets.csv") |>
dragon_map(prompt = "question", response = "answer") |>
dragon_train("Qwen/Qwen2.5-0.5B-Instruct", wait = TRUE)
dpo <- dragon_dataset("preferences.csv") |>
dragon_map_pairs(prompt = "question", chosen = "better", rejected = "worse") |>
dragon_prefer(sft, method = "dpo", beta = 0.1, wait = TRUE)
dragon_evaluate(dpo) # preference accuracy and reward margin on held-out pairs
dragon_generate(dpo, "My thermostat keeps dropping off Wi-Fi.")
Passing a run as the model chains the stages: the earlier adapters are
folded into the weights before the new stage adds its own. method = "orpo"
needs no reference model and can start from a base model directly. The app
has the same path: choose "Preference pairs" in the Map panel and a run to
start from in the Train panel.
When the goal is verifiable, a correct number, valid JSON, a format, a length budget, reinforcement learning beats preference data. The model writes several answers per prompt, the rewards score them, and it learns from the ones that beat their group's average (GRPO):
math <- dragon_dataset("arithmetic.csv") |>
dragon_map_prompts(prompt = "question", reference = "answer")
rl <- dragon_reinforce(
math, sft,
rewards = list(dragon_reward("numeric"), dragon_reward("length", max_chars = 300, weight = 0.2)),
group_size = 6, wait = TRUE
)
dragon_evaluate(rl) # mean held-out reward, per reward
Built-in rewards cover exact and numeric answers, regex and JSON formats,
length, and keywords; a "custom" reward points at a Python function that
sees the prompt, the completion, the reference, and the row's other
columns. It is the right tool for verifiable goals and the wrong one for
vague ones; use dragon_prefer() for "be more helpful".
Small models are only as good as their training data, and most teams do not have a few hundred hand-written ideal replies. Two shortcuts:
# A stronger model answers your prompts; its replies become the training set.
teacher <- dragon_llm_anthropic(system = "You are a concise, warm support agent.")
synth <- dragon_synthesize(dragon_prompts(sft, "train"), teacher, system = "You are a concise, warm support agent.")
sft2 <- dragon_train(synth, "Qwen/Qwen2.5-0.5B-Instruct", wait = TRUE)
# The run answers each prompt four times, a judge scores every sample, and the
# best and worst become preference pairs. Then DPO on top of the same run.
pairs <- dragon_synthesize_pairs(dragon_prompts(sft2, "train", n = 200), student = sft2,
judge = dragon_judge_anthropic(model = "claude-sonnet-5"))
dpo <- dragon_prefer(pairs, sft2, wait = TRUE)
dragon_judge(dpo, against = "base", judge = dragon_judge_anthropic())
That last sequence, sample, judge, train, judge again, is the loop that turns a fine-tune into a development cycle. The app's Try it panel has the same Improve step.
Training machines are rarely the right inference machines. Generation goes through a backend you choose:
chat <- dragon_chat(run, system = "You are a concise support agent.")
chat$say("My thermostat keeps dropping off Wi-Fi.")
chat$say("I tried that. What else?") # the model remembers the first exchange
backend <- dragon_serve_ollama(run) # merge, register with Ollama, done
options(dragonfarm.backend = backend) # every call now uses Ollama
dragon_generate(run, "Hello")
vllm <- dragon_backend_server("https://my-pod.example.com/v1", model = "me/support-0.5b")
dragon_chat(backend = vllm)$say("Hello")
The default backend is a local worker that keeps the last two models loaded, so judge and synthesis loops stop paying a model load per call. The app's Chat panel offers the same choice of backend.
Conversations are also where human feedback comes from. Rate a reply, fix
it, and the verdict is saved; dragon_feedback() turns the verdicts into
data for the next stage:
chat$rate("down")
chat$regenerate()
chat$rate("up")
chat$edit("Hold the recessed button on the back for ten seconds.")
fb <- dragon_feedback()
better <- dragon_train(fb$sft, dpo, wait = TRUE) # liked and edited replies, with context
dpo2 <- dragon_prefer(fb$pairs, better, wait = TRUE) # liked versus disliked replies to the same prompt
Whole conversations work as training data too: dragon_conversations()
takes a list of message lists or a JSONL file of them.
p <- dragon_pipeline("Qwen/Qwen2.5-0.5B-Instruct", list(
dragon_step_train(tickets),
dragon_step_synthesize_pairs(prompts = "train", n = 150, judge = dragon_judge_anthropic(model = "claude-sonnet-5")),
dragon_step_prefer(method = "dpo"),
dragon_step_judge(against = "base", judge = dragon_judge_anthropic()),
dragon_step_evaluate(metrics = c("token_f1", "length_ratio"))
), background = TRUE)
dragon_pipeline_status(p) # step by step, while it runs
dragon_compare(p) # its runs side by side, when it is done
The app's Pipeline panel runs the same recipe and draws the lineage of every run in the directory.
Held-out loss says a stage trained. It does not say the replies got better. Three tools answer that:
# Deterministic checks over every held-out row
dragon_evaluate(dpo, metrics = c("exact", "token_f1", "json_valid"))
# A stronger model as judge: did DPO beat the fine-tuned run it started from?
dragon_judge(dpo, against = "base", judge = dragon_judge_anthropic())
#> dpo vs sft on 20 prompts: wins 65% · ties 25% · losses 10%
# Or absolute scores against your own rubric, with a local judge
dragon_judge(sft, judge = "Qwen/Qwen2.5-1.5B-Instruct",
rubric = "Reward concrete next steps; penalise anything over 120 words.")
# Everything the package knows about every run, side by side
dragon_compare()
Pairwise judging asks each question twice with the replies swapped, so a
judge that favours whichever answer comes first yields ties, not wins.
dragon_judge_anthropic() reads ANTHROPIC_API_KEY; dragon_judge_ellmer()
accepts any ellmer chat for other providers.
The run directory is the whole contract between R and the trainer, so a run
can be trained on any machine with a GPU and its results copied back.
dragon_bundle() zips the run, dragon_remote() opens a provider with the
dragon-farm notebook and prints the steps, and dragon_import() puts the
trained adapter into place. Nothing else changes: dragon_generate() and
dragon_merge() work on the imported run as if it had trained locally.
run <- dragon_dataset(dragon_example_data()) |>
dragon_map(prompt = "{subject}\n\n{body}", response = "reply") |>
dragon_bundle("Qwen/Qwen2.5-0.5B-Instruct")
dragon_remote(run, "colab") # opens Colab with the notebook, prints the steps
# ... upload the zip it names, Run all, download dragonfarm-results-<id>.zip ...
dragon_import(run, "~/Downloads/dragonfarm-results-<id>.zip")
| Provider | Cost | What the link opens |
|---|---|---|
| Google Colab | Free tier with a T4; paid tiers for longer sessions | The notebook, directly |
| Kaggle | Free: about 30 GPU hours a week (T4 x2 or P100) | The notebook, directly |
| Lightning AI | Free monthly credits, then pay as you go | The dragon-farm repo in a new Studio |
| RunPod | Pay per hour, wide choice of GPUs | The RunPod console |
The app has the same path: the Train panel's "No GPU here?" section prepares
the bundle and gives you the download and the provider link, and the Monitor
panel imports the results zip. dragon_check() points here when it finds no
GPU, and a run that failed locally for lack of memory can be sent to the
cloud as is with dragon_remote(run, ...).
| File | Written by | Contents |
|---|---|---|
config.json |
R | Everything the trainer needs. |
data/train.jsonl, data/eval.jsonl |
R | Rows in chat format. |
status.json |
Python | State, device, parameter counts, final metrics. |
progress.jsonl |
Python | Loss, learning rate, and ETA per logging step. |
adapter/ |
Python | The LoRA adapter, loadable with peft. |
checkpoints/ |
Python | The last two checkpoints, for resume. |
eval.json, samples.json |
Python | Held-out loss and sample generations. |
merged/ |
Python | After dragon_merge(): a standalone model. |
| Model | Size | License | Needs a token | Min GPU memory |
|---|---|---|---|---|
HuggingFaceTB/SmolLM2-135M-Instruct |
135M | Apache 2.0 | no | 2 GB, or CPU |
HuggingFaceTB/SmolLM2-360M-Instruct |
360M | Apache 2.0 | no | 3 GB |
Qwen/Qwen2.5-0.5B-Instruct |
0.5B | Apache 2.0 | no | 3 GB |
google/gemma-3-1b-it |
1B | Gemma | yes | 5 GB |
meta-llama/Llama-3.2-1B-Instruct |
1.2B | Llama 3.2 | yes | 5 GB |
Qwen/Qwen2.5-1.5B-Instruct |
1.5B | Apache 2.0 | no | 7 GB |
HuggingFaceTB/SmolLM2-1.7B-Instruct |
1.7B | Apache 2.0 | no | 8 GB |
dragon_presets() returns this table. Any other causal language model on the
Hugging Face Hub works too. For gated models, accept the license on the Hub
and set HF_TOKEN in the R session.
R never imports torch. It writes a run directory and launches
python -m dragonfarm.train as a subprocess with processx. The trainer is
Hugging Face transformers with peft for LoRA and a prompt-masking collator
so only the reply tokens contribute to the loss. Progress comes back through
files, which is what lets the Shiny app poll it and lets a run outlive the R
session.
The Python side (inst/python) is also its own installable
package (pip install ./inst/python, soon pip install dragonfarm once it's
on PyPI): a dragonfarm command (check, train, generate, pack) and a
dragonfarm.api module for reading a run directory from plain Python, no R
required. A run started from R can be inspected or continued from Python and
back, since both read and write the exact same files.
devtools::test() # unit tests, no Python needed
Sys.setenv(DRAGONFARM_INTEGRATION = "true")
devtools::test(filter = "integration") # trains SmolLM2-135M for 6 steps
The Python side has its own tests: PYTHONPATH=inst/python python inst/python/tests/run.py (or python -m pytest inst/python/tests if pytest is installed).
| Variable | Effect |
|---|---|
DRAGONFARM_PYTHON |
Use this interpreter instead of the one reticulate builds. It must already have the packages from dragon_python_requirements(). |
DRAGONFARM_TORCH_INDEX |
Windows only. auto (default) selects the CUDA wheel index matching your NVIDIA driver on first Python use. Set to "" to use PyPI's CPU build, or to another index URL. |
DRAGONFARM_RUNS_DIR |
Where runs are stored. Defaults to dragonfarm_runs under a session temp directory; see Requirements for a persistent location. |
HF_TOKEN |
Hugging Face token for gated models. |
LLAMA_CPP_DIR |
A llama.cpp checkout, for dragon_export_gguf(). |
MIT.