- Reinforcement fine-tuning (RFT) trains a model by sampling answers, scoring them with a grader and reinforcing what scored well. It needs tens to hundreds of well-chosen tasks and a reliable way to score outcomes, rather than thousands of written-out ideal answers, which makes it the right tool where correctness can be checked but the path to it is hard to demonstrate.
- Since late 2025 RFT has been a managed service on the major platforms, including for reasoning models and for agents that call tools during training. The hard part moved from infrastructure to the grader: whatever the grader rewards is what you get, including its blind spots.
- Continual learning applies the same idea to production: traces, outcomes and corrections become training signal. Most enterprises should start with the cheap end, memory, skills and prompt updates gated by evaluations, and move to weight updates only for high-volume tasks with trustworthy graders and strict regression gates.
Learning from outcomes, not examples
The adaptation techniques in the adaptation ladder share an assumption: someone can show the model what good looks like. Supervised fine-tuning trains on example inputs paired with ideal outputs; distillation uses a stronger model's outputs as the examples. That works beautifully for format, style and narrow skills, and it breaks down where the ideal output is hard to write but easy to check. A correct medical-coding decision, a reconciliation that balances, a SQL query that returns the right rows, a contract clause classified the way legal would classify it: in each case an expert can verify an answer far more cheaply than they can author a perfect worked example.
Reinforcement fine-tuning exploits that asymmetry. Instead of imitating examples, the model is trained on outcomes: it attempts a task, a grader scores the attempt, and training shifts the model toward the attempts that scored well. This is the same family of techniques that produced reasoning models, where reinforcement learning against verifiable rewards, such as correct math answers and passing tests, taught models to think at length before answering, as the training pipeline deep-dive describes. What changed for enterprises is access. OpenAI offers reinforcement fine-tuning of a reasoning model with custom graders, and an agent variant that trains through live tool calls; Amazon Bedrock launched the capability in December 2025 and in February 2026 extended it to open-weight models. What used to need a research team is now an API call plus a grader.
How reinforcement fine-tuning works
The loop is simple to describe. Take a set of task prompts. For each, the model samples several candidate responses. A grader scores each candidate. The training step increases the probability of the responses that scored above the group's average and decreases the rest; group-relative methods popularized by DeepSeek's open research made this efficient without a separate value model. Repeat, and measure on held-out tasks after every round.
TASK SET (tens to hundreds of real, representative prompts)
|
v
MODEL samples k candidates per prompt
|
v
GRADER scores each candidate
exact match | unit tests | rule checks | rubric judge
|
v
UPDATE: reinforce above-average candidates,
discourage below-average ones
|
v
EVAL GATE on held-out tasks --fail--> stop, inspect grader
|
pass
|
v
next round, or ship the checkpoint
Three properties follow. RFT is data-efficient in examples but expensive in compute, because every training step samples several full responses and grades them. It improves the model on the distribution of tasks you train on and can degrade it elsewhere, so the held-out evaluation must cover what you care about, not only the target task. And because the model explores, it can find solutions nobody demonstrated, which is the point, and strategies that satisfy the grader without solving the problem, which is the risk.
The grader is the product
In supervised fine-tuning the dataset is the product. In reinforcement fine-tuning the grader is. Graders come in a ladder that mirrors the sensors in the harness engineering deep-dive. Deterministic graders check exact answers, run unit tests, validate schemas or apply business rules; they are cheap and hard to fool, and should be used wherever the task allows. Model-based graders score against a rubric for qualities rules cannot capture; they extend RFT to subjective tasks but inherit every weakness of an LLM judge, as the evaluation methods deep-dive explains. Many real graders combine both: a rule check that gates, followed by a judge that ranks.
Grader quality is measurable and should be measured. Calibrate it against expert judgment on a sample, track agreement, look for systematic leniency, and version it like code, because changing the grader changes what every future checkpoint learns.
When RFT is worth it
RFT is the most powerful rung on the adaptation ladder and the most demanding, so the decision deserves discipline:
| Approach | Needs | Best for | Avoid when |
|---|---|---|---|
| Prompting, skills, retrieval | Instructions and context | Most tasks; knowledge that changes | Rarely; always try first |
| Supervised fine-tuning | Hundreds to thousands of ideal examples | Format, style, narrow imitation | Ideal outputs are hard to write |
| Distillation | A stronger teacher model | Moving a capability to a cheaper model | The teacher also fails the task |
| Reinforcement fine-tuning | Representative tasks and a trustworthy grader | Checkable decisions where reasoning matters | Quality cannot be scored reliably |
Three tests settle most cases. Can correctness be checked automatically, or by a calibrated judge, at reasonable cost? Is the task high-volume or high-value enough to repay the training compute and the ongoing re-training when base models change, the treadmill described in the fine-tuning in practice deep-dive? And has a strong prompted baseline with good skills and retrieval actually been tried and measured? If any answer is no, RFT is premature.
Running an RFT project
A first reinforcement fine-tuning project fails most often for reasons that have nothing to do with training. A sequence that avoids the common traps:
- Pick a narrow, checkable task. One decision type with a clear notion of correct, high enough volume to matter, and a current error rate that hurts. "Improve the support assistant" is not a task; "assign the right refund reason code" is.
- Build the task set from production. Tens to a few hundred real cases, deliberately including the hard and ambiguous ones, plus a separate held-out set the training never touches. Synthetic tasks, covered in the synthetic data deep-dive, can widen coverage but should not replace real cases.
- Establish the baseline. Measure the best prompted configuration, with skills and retrieval, on the held-out set. This is the number RFT has to beat, and often the project ends here because the baseline turns out to be good enough.
- Build and calibrate the grader. Score a sample with experts, compare, fix disagreements, and attack the grader for loopholes. Budget more time here than for training.
- Run small, then read the outputs. A short run on a subset shows whether scores move and, more importantly, how. Read samples from each round; rising scores with strange outputs mean the grader is being gamed.
- Gate, deploy, and plan the re-run. Promote a checkpoint only if it beats the baseline on the held-out set without regressions elsewhere, route it through the gateway like any model version, and schedule the re-training that every base-model upgrade will require.
Cost is worth estimating before step five. Training compute scales with the number of tasks, the number of samples per task and the length of each response, and reasoning models produce long responses. The ongoing cost is usually larger than the first run: re-evaluation and re-training on each new base model, grader maintenance as the business changes, and the inference premium or savings of the tuned model. If a smaller tuned model replaces a larger general one, the inference savings can pay for all of it, which is the most common business case.
Continual learning in production
The same logic, applied to a live system, is continual learning: production traces, outcomes and corrections become the signal that improves the next version. In 2026 it became a practical concern rather than a research topic. Coding-tool vendors reported reinforcement learning on production interactions with updates shipped every few hours, and always-on agents advertised that they learn from feedback over time.
For most enterprises the right place to start is not the weights. Continual learning has a cheap end and an expensive end. At the cheap end, the system improves its memory, skills, prompts, retrieval and routing from what production reveals: a recurring failure becomes a new guide or check, a correction becomes a remembered preference, a frequent question becomes a verified answer. These changes are fast, reversible and inspectable, and they compound, as the agent memory deep-dive describes. At the expensive end, production data feeds periodic reinforcement or supervised fine-tuning of the model itself. That captures improvements no prompt can, but it is slower to verify, harder to reverse and subject to every data-governance question about using customer interactions for training.
Either way the process is the same: collect traces and outcomes, find a recurring failure, propose a change, prove it improves the target without unacceptable regression elsewhere, then promote it. The evaluation gate is what separates continual learning from continual drift.
Reward hacking and other failure modes
Reinforcement learning's characteristic failure is reward hacking: the model finds a way to score well without doing the task, such as special-casing the tests, exploiting a judge's preference for confident tone, or formatting answers to slip past a rule check. Defenses are layered graders, held-out evaluations the training never sees, human review of samples from each round, and monitoring for scores that rise faster than real-world outcomes. Other failures are more mundane: forgetting capabilities outside the training distribution, overfitting to a small task set, feedback loops where the model learns from its own earlier mistakes recorded as outcomes, and privacy exposure when production data enters training. Each has a control, and each control is a reason to start at the cheap end of continual learning.
The architect view
Reinforcement fine-tuning changes the economics of adaptation for one specific class of problem: decisions that are expensive to demonstrate but cheap to check. For those, a few hundred representative tasks and a solid grader can buy improvements that no amount of prompt work delivers, and managed services have removed most of the infrastructure barrier.
Four commitments keep it disciplined. Build graders as first-class, versioned assets, and red-team them before training. Prove a strong prompted baseline before training anything. Gate every checkpoint, and every learned change in production, on held-out evaluations that include what you must not break. And treat continual learning as a ladder: memory, skills and prompts first, weights only where volume and grader quality justify it.
The deeper shift is that evaluation assets become training assets. The organization that has invested in good evals and good graders can now turn them directly into better models; the one that has not will find it cannot adapt models it cannot measure.