Skip to content

T01 · Training Plans Are Not an Algorithm Leaderboard: Layer the Decision First ​

In one sentence: Training methods should not be ranked by how new their technical names sound. The real decision order is to examine the capability goal and task shape first, then the data and learning signal, then the assumptions each candidate method depends on, and finally how it changes model components, the training system, the resource bill, and failure modes.


🧠 Foundation Model Training Decisions, Chapter 1 · One Skill to Practice

This track assumes that the team already operates a foundation-model production system; it does not discuss whether an ordinary team should train a large model. This chapter first uses a concrete example to explain SFT, DPO, PPO, GRPO, and related concepts, then establishes a common coordinate system for later chapters. You do not need to derive formulas after reading it, but you should be able to understand what data a training method needs, which components it adds, what problem it solves, and where it moves the complexity.


Reading Map: Where This Chapter Sits in Model Production ​

Before the example, place this chapter back into the full lifecycle:

text
Large-scale raw data
    ↓ Pre-training: repeatedly predict the next token
Base Model
    ↓ Post-training: demonstrations, preferences, rewards, tools, and safety training
Candidate product model
    ↓ Capability / safety / regression evaluation
Released model
  • Pre-training: enables a model to learn language, knowledge, and basic capabilities from large-scale data; its main objective is usually next-token prediction.
  • Base Model: the model produced by pre-training. It can already continue text and solve some problems, but it may not follow instructions reliably or have the behavior and safety boundaries a product requires.
  • Post-training: continues training the Base Model so that it follows instructions, expresses desired preferences, reasons, uses tools, and respects safety constraints more effectively.
  • Fine-tuning: continues training from an existing model's parameters instead of training from scratch. SFT, DPO, and reinforcement-learning post-training may all change existing model parameters.

Here, parameters (or weights) can initially be understood as the enormous collection of learnable numbers inside a model. Training does not write individual if rules into the model. It adjusts these numbers so that the tokens, answers, or actions we want become more probable.

This chapter focuses on post-training because PPO, GRPO, DPO, and RLHF mainly appear there. It ends, however, with a coordinate system that can also analyze pre-training decisions.

The minimum distinction: supervised learning mainly teaches a model to imitate existing correct demonstrations. Reinforcement learning lets the model attempt the task, uses Reward to judge the result, and then adjusts the Policy so that high-return behavior appears more often.


Opening: A Minimal Post-training Scenario ​

Suppose we want to improve a model's mathematical ability and give it this problem:

Calculate 17 × 23 and show brief steps.

The model makes four attempts:

AttemptModel responseProgram check
A17 × 23 = 3911 point
B17 × 23 = 3810 points
C20 × 23 - 3 × 23 = 3911 point
D17 × 23 = 4010 points

This small table already contains the most important objects in large-model reinforcement learning:

text
Prompt (the problem)
    ↓
Policy / Actor (the model currently being trained)
    ↓ Generate four times
Rollout (let the model actually try)
    ↓
Response / Trajectory (the answer or complete action trace)
    ↓
Verifier (the checking program)
    ↓
Reward (0 or 1 point for each attempt)
    ↓
Turn "which behaviors should occur more often" into parameter updates

The hard part is the last step. A and C earned 1 point, so we naturally want the model to generate them more often; B and D earned 0, so we want them less often. But by exactly how much should each probability rise or fall? If the task is not a one-step response but a Coding Agent that works for two hours and invokes dozens of tools, which intermediate step deserves credit or blame for the final success or failure?

Methods such as PPO and GRPO mainly address this kind of problem. Before comparing them, however, we must identify every role in the diagram.


1. Identify the Basic Roles in a Training System ​

1.1 Prompt, Response, and Token ​

  • Prompt: the input given to the model. It may be one problem, or it may include system instructions, a user request, tool descriptions, and prior conversation.
  • Response / Completion: the output the model generates for the Prompt.
  • Token: the basic unit a model uses to read and write text. A Chinese character, part of an English word, or a punctuation mark may occupy one or more Tokens.

A model does not write an entire response at once. It repeatedly calculates what the next Token should be and generates the response one Token at a time.

1.2 Policy and Actor: The Model Currently Making Decisions ​

A Policy can initially be understood as:

The complete set of probability rules a model uses to decide “what to generate next” given the current input.

For the same Prompt, a model might assign 391 a probability of 40%, 381 a probability of 10%, and all other answers a combined probability of 50%. Training is, at its core, the process of adjusting these probabilities.

In a reinforcement-learning system, when this model actually generates answers or performs actions, it is often called the Actor. In many large-model training systems, Policy Model and Actor refer to different roles of the same model:

  • Policy emphasizes that it is the strategy being optimized;
  • Actor emphasizes that it executes the strategy in the Rollout system and produces trajectories.

Do not assume that the Actor must be a separate new model.

1.3 Rollout and Trajectory: The Attempt and the Trace It Leaves ​

A Rollout is the process of letting the current model actually attempt a task once.

  • For a math problem, one Rollout may simply generate one complete answer.
  • For a Coding Agent, one Rollout may include reading code, searching, editing files, running tests, and continuing to repair the result until it succeeds, fails, or reaches a step limit.

The complete record left after a Rollout is called a Trajectory. It usually includes:

text
Input → model output → tool call → environment result → next output → … → final result

Papers and engineering documents sometimes use Rollout and Trajectory interchangeably. This track uses the following convention:

  • Rollout emphasizes the act of generating one unit of experience;
  • Trajectory emphasizes the training data saved after generation.

1.4 Reward, Verifier, and Reward Model: Who Assigns the Score ​

A Reward is the score assigned to the result of an answer or action.

Someone—or something—must calculate that score. Two common types of scorer are:

ScorerHow it scoresSuitable forMain risk
VerifierDirectly checks rules, answers, compiler output, unit tests, or environment stateVerifiable tasks such as math, code, formatting, and tool useThe model may exploit loopholes, peek at tests, or take shortcuts
Reward ModelLearns “which response is better” from human or AI preference data, then outputs a scoreHelpfulness, tone, safety, and other concerns that are hard to encode as rulesThe Reward Model may be biased or exploited by the Policy

Keep these distinctions clear:

  • Reward is the final score;
  • a Verifier or Reward Model is the component that produces the score;
  • a Critic is not the Reward Model that grades the result.

1.5 Value and Critic: Estimating the Future Score Before the Task Ends ​

Suppose a Coding Agent has completed 20 steps but the task is not finished. We want to know:

From the current state, how much final reward can we expect if it continues?

This expectation is called Value. The model that learns and estimates Value is called the Critic, or value model.

The difference between a Reward Model and a Critic is easy to remember:

ComponentCore question
Reward Model / Verifier“How good is the result that has already been produced?”
Critic“From the current state, how good can the future result be expected to become?”

A Critic can provide a baseline before a task ends, helping us judge whether a step performed better than expected. The cost is that the team must train and run an additional model, and the Critic itself may be wrong.

1.6 Advantage: How Much Better the Actual Result Was Than Expected ​

Advantage answers:

Compared with a baseline, how much better or worse was this action's result?

For now, remember this intuitive relationship:

text
Advantage ≈ actual return - expected return

For example, suppose a Critic estimates that the current probability of solving a problem is only 0.35:

  • a correct answer receives 1 point, so performance exceeded expectations and Advantage is positive;
  • an incorrect answer receives 0 points, so performance fell below expectations and Advantage is negative.

Training generally increases the probability of positive-Advantage behavior and reduces the probability of negative-Advantage behavior. “Updating the model” means adjusting model parameters through optimization. Real algorithms also consider discounting, per-Token attribution, and update stability, but this intuition is enough for Chapter 1.

1.7 Credit Assignment: Which Step Deserves Credit or Blame ​

Credit Assignment is not about handing out awards. It asks:

How should the final score of 1 or 0 be distributed across every intermediate Turn, Action, or Token?

A one-step math problem is simple; a long-horizon Agent is difficult:

  • it selects the wrong file at step 3, but the test fails only at step 40;
  • it discovers critical information at step 10, but completes the task only at step 80;
  • a step appears to fail, yet rules out a wrong path for everything that follows.

Giving the entire trajectory one final score produces a cheap but coarse signal. Assigning different Advantage values to every step or even every Token gives a finer signal, but requires a Critic, process verifier, or another estimation mechanism.

1.8 Reference Policy: Preventing the Model from Moving Too Far at Once ​

A Reference Policy is usually a frozen model snapshot from before training begins. It does not score results. Its main purpose is to answer:

How far has the new model's behavior moved from the original model?

DPO compares the current Policy with the Reference Policy in terms of their relative preferences for answers. Many PPO-style large-model post-training systems also use the Reference Policy to calculate a KL penalty, preventing the Policy from drifting sharply while chasing Reward.

Here, KL can initially be understood as “the distance between two probability distributions”; no derivation is needed in this chapter.

1.9 On-policy, Policy Lag, and Compaction ​

These three terms appear in the real-world case later:

  • On-policy: training data is generated by the current Policy, or a version very close to it. Ideally, each model update is followed by resampling with the new model.
  • Policy Lag: Rollout still generates data with an older version while the Trainer has already advanced through several versions. Asynchronous systems gain throughput, but their data becomes increasingly stale.
  • Compaction: when the context of a long-horizon Agent grows too long, earlier history is compressed into a summary or new state representation before the task continues. During training, this can split one very long trajectory into multiple sub-trajectories with different counts and lengths.

These are not incidental engineering terms beside the algorithm. They change the shape of the training data and therefore determine whether a method's assumptions still hold.

1.10 Online and Offline: Must the Model Keep Trying During Training? ​

  • Offline Training: the data is prepared before training begins. Training mainly rereads fixed data, such as SFT demonstrations or DPO preference pairs.
  • Online Rollout: during training, the current Policy continuously generates new answers or interacts with an environment, and the new experience is sent back to the Trainer.

“Online” does not mean that the model is publicly available. It means that data generation and model updates form a loop. Online methods can explore new behavior produced by the current Policy, but generating data is expensive, and the system must handle Policy Lag, environment failures, and coordination between training and inference resources.


2. How Four Common Methods Actually Teach a Model ​

With the roles above in place, we can examine SFT, DPO, PPO, and GRPO. For clarity, all four methods continue using the example of teaching a model to solve a math problem.

2.1 SFT: Show the Model a Good Answer Directly ​

The data required by SFT (Supervised Fine-Tuning) is the most intuitive:

text
Prompt: Calculate 17 × 23 and show the steps.
Demonstration: 20 × 23 - 3 × 23 = 460 - 69 = 391.

The training objective makes the model more likely to generate the Tokens in the demonstration after seeing the Prompt.

text
Prompt + high-quality demonstration
        ↓
Directly train the Policy to imitate the demonstration
        ↓
New Policy

SFT is suitable when:

  • clear, high-quality demonstrations can be written;
  • the goal is to teach instruction formats, response style, tool protocols, or basic task procedures;
  • the team wants to establish target behavior through the simplest and most stable route first.

Its limitation is that the model mainly imitates existing demonstrations. Errors not covered by the demonstrations, complex environment interaction, and questions such as “both responses answer the question, but which is better?” are difficult to solve through SFT alone.

In one sentence: SFT answers “What does a correct demonstration look like?”

2.2 DPO: Do Not Supply One Correct Answer; Say Which One Is Preferred ​

Some questions have no single standard answer, but people can judge which of two responses is better. We can prepare preference pairs:

text
Prompt: Explain why 17 × 23 = 391.

Chosen (better): 20 × 23 - 3 × 23 = 460 - 69 = 391.
Rejected (worse): The result is 391. Trust me.

DPO (Direct Preference Optimization) directly trains the Policy to prefer the Chosen response and disfavor the Rejected response relative to a Reference Policy.

text
Prompt + Chosen + Rejected
            │
Reference Policy (frozen)
            ↓
Directly update the current Policy

The core DPO training stage usually does not require:

  • continuously generating new online Rollouts with the Policy;
  • training a separate Reward Model;
  • running a Critic.

The system therefore resembles an offline preference-data training pipeline and is usually simpler than online RL. The trade-off is that the model mainly learns within the coverage of the existing preference data. If the data contains no example of a new strategy, the model will also struggle to discover it through online exploration. Original DPO paper

In one sentence: DPO answers “Of these two existing answers, which should the model favor?”

2.3 Critic-based PPO: Attempt the Task, Receive Reward, Then Make a Constrained Update ​

PPO (Proximal Policy Optimization) is a family of policy-update methods. Proximal can be understood as “do not move too far at once”: even if a batch of Rollouts scores highly, PPO limits the size of a single parameter update so the Policy does not change too abruptly. Original PPO paper

An important distinction is:

PPO itself mainly specifies how to make constrained policy updates. It does not require Reward to come from humans, nor does its definition strictly require a Critic.

In online reinforcement learning for large models, however, Critic-based PPO is common: the Actor generates Rollouts, a Reward Model or Verifier scores them, and a Critic estimates Value to calculate finer-grained Advantage.

text
Current Actor / Policy
        ↓ Online Rollout
Answer or multi-turn trajectory
        ↓
Reward Model / Verifier assigns Reward
        ↓
Critic estimates Value → calculate Advantage
        ↓
PPO constrains the size of each update
        ↓
Update the Actor while avoiding excessive drift from the Reference Policy

It is worth considering when:

  • the model can interact repeatedly with the task environment;
  • trajectories are long and irregular in length;
  • even a single Rollout should provide training value;
  • Turn-level or Token-level fine-grained credit assignment is needed;
  • the team can afford the training and resource cost of a Critic.

Its system costs are equally direct:

  • the Actor must generate data;
  • the Critic must be trained and run for inference;
  • a Reward Model or Verifier must provide rewards;
  • a Reference Policy is often used to limit drift;
  • multiple models share or compete for GPUs, GPU memory, and network capacity;
  • Rollout and training versions must remain sufficiently close.

In one sentence: Critic-based PPO answers “After the model actually tries, which actions performed better than expected, and how should the Policy update without changing abruptly?”

2.4 GRPO: Compare Multiple Attempts on the Same Prompt Instead of Training a Critic ​

GRPO (Group Relative Policy Optimization) was introduced in DeepSeekMath and is a variant of PPO. Its key change is that it does not rely on a separate Critic. Instead, the same Prompt generates a group of responses, and their within-group scores form a relative baseline. DeepSeekMath paper

Return to the four opening attempts:

ResponseRewardPerformance relative to group average
A1Above group average; positive Advantage
B0Below group average; negative Advantage
C1Above group average; positive Advantage
D0Below group average; negative Advantage

The intuitive GRPO flow is:

text
The same Prompt
    ↓ Current Policy generates K responses
Response 1, response 2, …, response K
    ↓
Verifier / Reward Model scores each one
    ↓
Calculate relative Advantage from the within-group mean and distribution
    ↓
Update the Policy with a PPO-style constrained objective

Removing the Critic reduces the number of models, GPU memory use, and value-learning complexity, but introduces new prerequisites:

  • the same Prompt can generate multiple samples;
  • those samples are comparable;
  • Reward varies enough within the group;
  • generating K samples is affordable;
  • if environment failures cause missing samples, the system can refill the group or handle an incomplete group correctly.

If every response in a group receives 1, or every response receives 0, each result equals the group average and the relative signal degenerates. For very long Agent trajectories with widely varying lengths or Compaction, even the definition of “a group of samples” can become difficult.

In one sentence: GRPO answers “Across multiple attempts at the same problem, which attempts are better relative to the group?” It exchanges grouped sampling for the Critic.

2.5 Comparing the Four Methods ​

MethodMinimum training dataOnline Rollout requiredSeparate CriticPrimary benefitMain cost or limitation
SFTPrompt + demonstrated answerNoNoImitates explicit target behaviorDepends heavily on demonstration quality and coverage
DPO familyPrompt + Chosen + RejectedUsually not in the core training stageNoLearns preferences through a simpler offline pipelineLimited by offline data distribution and preference quality
GRPO familyPrompt + multiple same-prompt Rollouts + RewardYesUsually noLearns through group comparison without a CriticRequires groups that can be formed, compared, and exhibit Reward variance
Critic-based PPOOnline Rollout + Reward + Value estimateYesYesLearns from single trajectories with finer-grained credit assignmentHigher costs for the Critic, GPUs, synchronization, and stability

This table is not a ranking of sophistication. It simply shows that different input data and learning problems require different system components.


3. Why RLHF, RLAIF, and Verifiable Rewards Are Not Peers of PPO ​

Consider this commonly conflated list:

text
RLHF vs RLAIF vs DPO vs PPO vs GRPO

The problem is that these concepts belong to different decision layers.

3.1 RLHF: Feedback Mainly Comes from Humans ​

RLHF (Reinforcement Learning from Human Feedback) describes a family of processes that shape model behavior using human feedback.

The classic InstructGPT route is:

text
Human demonstrations → SFT
Humans compare model responses → train a Reward Model
Reward Model scores responses → PPO updates the Policy

PPO can therefore be the optimization method used inside an RLHF process. They are not mutually exclusive options. InstructGPT paper

3.2 RLAIF: Feedback Mainly Comes from AI ​

RLAIF (Reinforcement Learning from AI Feedback) mainly replaces direct human evaluation with AI evaluation to reduce labeling cost and increase data scale.

It still does not dictate whether PPO or another method must be used at the end. AI feedback can be used to:

  • produce preference pairs for DPO;
  • train a Reward Model and then run PPO;
  • combine with rules, safety principles, or human spot checks to form rewards.

3.3 Verifiable Rewards: Feedback Comes from Rules, Programs, or Environments ​

Mathematical answers, unit tests, compiler results, game outcomes, and tool-task completion states can all produce verifiable rewards. This describes the source of Reward, not a specific optimizer.

The same verifiable reward can work with:

  • GRPO: sample multiple results for the same Prompt and compare them within the group;
  • Critic-based PPO: learn from a single or long-horizon trajectory;
  • other policy-optimization methods.

The correct layering is therefore:

Decision layerQuestion it answersTypical options
Feedback sourceWho defines “better”?Humans, AI, rules, Verifier, environment
Data organizationWhat does the training material look like?Demonstrations, preference pairs, single Rollouts, grouped Rollouts, long-horizon trajectories
Credit assignmentHow is the final result attributed?Whole trajectory, process rewards, Critic, group-relative baseline
Policy updateHow do parameters change stably?SFT, DPO, PPO, GRPO, and others
System implementationWhich components are required?Actor, Critic, Reward Model, Verifier, Reference, Rollout Workers

Architecture insight: when asked “A or B?”, the first question is not which is stronger, but whether A and B answer the same question. If one describes where feedback comes from and the other describes how parameters update, they can coexist and do not belong in a multiple-choice question.


4. Distinguish Three Kinds of “Architecture” ​

“Pre-training architecture” and “post-training architecture” can be ambiguous because training discussions involve at least three structures.

4.1 Model Architecture: How One Forward Pass Happens ​

This describes the model's internal computation, for example:

  • Dense: every Token generally passes through the same dense parameters at each layer;
  • MoE: primarily in parts of the feed-forward network, only a small number of routed experts are activated;
  • Attention, which determines what contextual information should matter when generating the current Token;
  • multimodal encoders and connectors.

Model architecture mainly changes parameter capacity, activation computation, GPU memory, communication, and inference characteristics. A GPU is a parallel computing device commonly used to train models; GPU memory is the GPU's high-speed memory for parameters, activations, and other training data. This chapter does not expand on Dense or MoE. It is enough to know that they occupy a different layer from PPO and GRPO: the former change how a model computes; the latter change how it learns.

4.2 Training Method: Which Signals Update the Model ​

This describes the learning objective and update mechanism, for example:

  • Next-token Prediction;
  • SFT;
  • DPO;
  • PPO and GRPO;
  • distillation, where a stronger Teacher Model produces signals to train a smaller or cheaper Student Model.

It changes the form of training samples, the loss or reward signal, whether online Rollout is required, and how parameters are updated.

4.3 Training-System Architecture: Which Components Run the Method Together ​

This describes the engineering system, for example:

  • data production and version management;
  • Actor, Critic, Reward Model, and Verifier;
  • Rollout clusters and training clusters;
  • synchronous or asynchronous pipelines;
  • Checkpoints, evaluation gates, and model registries.

The three ultimately form a chain of constraints:

text
Capability goal
  ↓
Model architecture + training method
  ↓
Required data, computation graph, and auxiliary models
  ↓
Training-system topology, resource bill, and failure boundaries
  ↓
Whether evaluation evidence proves the choice worthwhile

Choosing GRPO is not merely changing a Loss—the computation rule a system uses to measure deviation from its training objective. The system no longer trains a Critic, but it must generate multiple comparable Rollouts for each Prompt and preserve complete sample groups. Choosing Critic-based PPO is likewise not merely swapping algorithms: the system must deploy and train a Critic, allocate GPU memory and GPUs to it, and handle bias in Value estimates.

Decision boundary: a technology becomes an architectural concern for this tutorial only when it changes the data shape, computation graph, system topology, resource bill, or failure modes.


5. An Eight-Layer Coordinate System for Training Decisions ​

Any future model-training proposal can be decomposed from top to bottom along these eight layers. Higher layers are closer to “why train”; lower layers are closer to “how the system supports it.”

text
① Capability goal: exactly what should improve?
        ↓
② Training material: what does the model actually see?
        ↓
③ Learning signal: how does the system define “better”?
        ↓
④ Signal granularity and credit assignment: which step receives credit or blame?
        ↓
⑤ Optimization method: how are parameters updated?
        ↓
⑥ Models and auxiliary components: which models are needed at runtime?
        ↓
⑦ Training-system topology: how are these components generated, trained, and synchronized?
        ↓
⑧ Promotion evidence: how do we prove the benefit exceeds the new cost?

① Capability Goal: Exactly What Should Improve? ​

Do not begin with “we need RL.” First define testable goals: math-answer accuracy, code-test pass rate, multi-turn tool-task completion rate, instruction following, or accurate safety refusals.

A code task that can be checked automatically and a subjective writing preference should not use the same reward-production mechanism.

② Training Material: What Does the Model Actually See? ​

The data may be raw Tokens, Prompts and demonstrated answers, Chosen/Rejected preference pairs, groups of candidate outputs for the same problem, or multi-turn trajectories containing environment feedback.

DPO requires offline preference pairs. Group-relative methods require multiple comparable samples for the same problem. Long-horizon Agent RL must handle environment state, tool results, and delayed rewards.

③ Learning Signal: How Does the System Define “Better”? ​

Signals may come from target Tokens, human preferences, AI judgments, rules, a Verifier, or the environment's final state.

Human evaluation is expensive and disputed. AI evaluation scales but inherits model bias. Verifiable rewards are cheap and explicit, but may tempt the model to exploit the Verifier.

④ Signal Granularity and Credit Assignment: Which Step Receives Credit or Blame? ​

Reward may apply only to an entire Sequence, or it may be refined to a Turn, Step, or Token.

Coarser granularity makes the signal easier to obtain but makes intermediate contributions harder to judge. Finer granularity makes attribution more expressive, but requires a Critic, process supervision, or reliable intermediate verifiers and introduces new estimation errors.

⑤ Optimization Method: How Are Parameters Updated? ​

Only at this layer should the method be selected:

  • if high-quality demonstrations exist, consider SFT first;
  • if offline preference pairs exist and online sampling is unaffordable, a DPO-style route is more natural;
  • if multiple comparable answers for one prompt can be generated cheaply, within-group Reward varies, and Critic cost matters, a GRPO-style route is more natural;
  • if trajectories are long and irregular, single samples must still be learned from, and fine-grained credit assignment is needed, Critic-based PPO is more natural.

“More natural” only means that a method's assumptions better match the problem. It does not remove the need for experimental validation.

⑥ Models and Auxiliary Components: Which Models Are Needed at Runtime? ​

MethodMain componentsCost assumed directly by the system
SFTPolicyDemonstration quality and coverage
DPO familyPolicy + Reference Policy, or precomputed reference Log-probs—the probabilities the reference model assigns to a passagePreference-pair production, reference computation, offline distribution limits
GRPO familyPolicy + Reward/Verifier, usually without a separate CriticMultiple samples per Prompt, group completeness, within-group variance
Critic-based PPOActor + Critic + Reward/Verifier, usually plus a Reference PolicyMulti-model GPU memory, GPU scheduling, Value learning, and synchronization

Removing a Critic is not free: while saving a model and GPUs, the system assumes the constraints of grouped sampling and relative comparison.

⑦ Training-System Topology: How Are Components Generated, Trained, and Synchronized? ​

text
Task / Prompt
    ↓
Rollout Workers (processes that generate experience)
    ├──────────────▶ Environment / tools / Verifier
    ↓
Trajectories and rewards
    ↓
Trainer (updates parameters) ◀── Policy / Critic / Reference
    ↓
New Policy version ──▶ Next round of Rollouts

At this point, determine whether Rollout and training are synchronous or asynchronous; whether the Actor, Critic, and Verifier are colocated or separated; how stale data from an old Policy may become before it is unusable; and whether environment failures should trigger sample replacement, discard the whole group, or retain partial trajectories.

A choice in an algorithm paper becomes a production problem involving queues, scheduling, versions, fault tolerance, and observability.

⑧ Promotion Evidence: How Do We Prove the Benefit Exceeds the New Cost? ​

Completing training does not prove the method is sound. At minimum, the evidence must show that:

  • the target capability improved;
  • non-target capabilities did not regress unacceptably;
  • training was sufficiently stable;
  • gains per unit of compute were reasonable;
  • Reward Hacking did not inflate scores;
  • the conclusion came from controlled ablations rather than simultaneous changes to data, algorithm, and model scale.

In particular:

  • Reward Hacking: the model exploits a scoring rule to obtain a higher score without improving real capability—for example, a Coding Agent reads hidden test answers instead of solving the problem.
  • Ablation Study: changes one factor at a time as far as possible while keeping the rest constant, revealing whether a gain came from data, reward, algorithm, or model scale.

Architecture insight: a training method does not move directly from a paper title to cluster configuration. It must pass through eight translation layers: goal → material → signal → attribution → optimization → components → system → evidence.


6. Real-world Case: Why GLM-5.2 Switched to Critic-based PPO ​

With the basic concepts in place, consider the chapter's real-world case.

The GLM-5 technical report explains that its RL algorithm was built on GRPO and that Agentic RL used group-based policy optimization: generating multiple trajectories for the same problem and deriving the training signal from within-group relative outcomes.

With GLM-5.2, the task shape changed. The official explanation describes the training problem for extremely long-horizon Agent tasks:

text
Tasks become long-horizon, multi-turn Agent interactions
        ↓
Very long trajectories undergo Compaction and are split into trainable sub-trajectories
        ↓
The same Prompt produces different counts and widely varying lengths of sub-trajectories
        ↓
The assumption of sampling one regular, comparable group of trajectories per Prompt no longer fits naturally
        ↓
Switch to Critic-based PPO
        ↓
Learn from a single Rollout, with the Critic estimating Token-level Advantage

Map the case back to the eight-layer coordinate system:

Decision layerGroup-based route in the GLM-5 reportGLM-5.2 long-horizon constraint
Capability goalAgentic, Reasoning, CodingLonger and more complex Agentic engineering tasks
Training materialGroups of multiple trajectories generated for one PromptSub-trajectories with irregular counts and lengths after Compaction
Credit assignmentGroup-normalized relative AdvantageToken-level Advantage estimated by a Critic
OptimizationGRPO and group-based policy optimizationCritic-based PPO and single-Rollout training
Model componentsNo separate Critic requiredAdd a Critic
System impactMust generate and preserve comparable sample groupsAdditional resources and synchronization for Actor, Critic, and Rollout
Main risksAll-correct/all-wrong groups, incomplete groups, incomparable trajectoriesCritic bias, unstable Value, higher resource cost

The official slime documentation illustrates the system cost directly: GRPO needs no Critic, while its PPO workflow must allocate GPU resources separately for the Actor, Critic, and Rollout.

Therefore, GLM-5.2's decision was not:

PPO is more advanced than GRPO.

It was:

Long, variable-length trajectories that have undergone Compaction no longer satisfy the prerequisites for regular within-group comparison. The team therefore accepted the Critic's training and resource cost in exchange for single-trajectory learning and Token-level credit assignment.

This does not mean Zhipu abandoned group-relative methods for every task. If answers are short, multiple comparable results can be generated cheaply for the same problem, Reward has enough within-group variance, and the Critic is expensive, a GRPO-style route may still be more reasonable.

The reusable lesson is not the answer “choose PPO,” but this chain:

text
Task shape
  → Data and trajectory shape
  → Assumptions required by the method
  → Model and system topology
  → New costs and failure modes
  → Promotion evidence

7. Applying the Same Coordinate System to Pre-training and Post-training ​

Pre-training and post-training usually retain the main backbone of the same model family, but their data, learning signals, and control systems differ.

Decision layerTypical question in pre-trainingTypical question in post-training
Capability goalGeneral language, code, math, multilingual, and long-context foundationsInstruction following, preferences, reasoning, tool use, safety behavior
Training materialLarge-scale Token sequences and data mixturesDemonstrations, preference pairs, synthetic data, Rollout trajectories
Learning signalMainly target TokensDemonstration targets, preferences, Reward, Verifier, environment feedback
Signal granularityUsually a Token-level training objectiveSequence-, Turn-, Step-, or Token-level feedback and attribution
Optimization choiceCoordination of data recipe, model structure, optimizer, and parallel trainingCombinations of SFT, DPO, PPO, GRPO, distillation, and others
Auxiliary componentsData pipeline, Trainer, Checkpoint, foundational evaluationReward Model, Critic, Verifier, environment, Rollout Workers
System bottleneckData throughput, GPU memory, communication, numerical stability, failure recoveryRollout throughput, reward quality, multi-model resources, Policy Lag
Promotion evidenceFoundational capabilities, gains from scaling data/model/compute, contamination checks, training stabilityTarget behavior, capability regressions, safety, Reward Hacking, real-task performance

Later chapters can explore Dense/MoE, data recipes, parallelism strategies, Checkpoints, SFT, preference optimization, and reinforcement learning separately, but each chapter will continue using this common decision language.


8. Write Every Training Decision as an ADR ​

Training experiments often record configurations and final scores but omit why a route was chosen. The ADR (Architecture Decision Record) method from 08 · Architecture Decision Records and Evolution can be applied directly to model training.

markdown
# ADR-Txx: Training-method decision title

## 1. Capability Goal
What should improve? Which metrics decide? Which capabilities must not regress?

## 2. Workload and Training Material
Are samples Tokens, preference pairs, or interactive trajectories? Single-turn or long-horizon? How are length and quality distributed?

## 3. Learning Signal and Credit Assignment
Who defines “better”? Does the signal apply to a Sequence, Turn, Step, or Token?

## 4. Candidate Methods and Their Assumptions
What data, sampling, model, and stability prerequisites does each method require?

## 5. Decision
What do we choose, and why does it better match the current constraints?

## 6. System Impact
Which models, data pipelines, queues, GPU roles, Checkpoints, and monitors are added or removed?

## 7. Costs and Failure Modes
Where has the complexity moved? What is the most likely failure?

## 8. Validation and Promotion Gates
Which ablations will we run? Which metrics must improve? Which regressions block Checkpoint promotion?

## 9. Exit Conditions
Which signals trigger a method change or rollback?

This template forces a team to translate “a paper reported strong results” into assumptions that its own system can test.


9. Five Common Decision Errors ​

Error 1: Ranking Sophistication by Publication Date ​

A new method may be more suitable only under specific data, task, model, or hardware constraints. The conclusion does not automatically transfer beyond those constraints.

Error 2: Turning Concepts from Different Layers into Mutually Exclusive Choices ​

“RLHF or PPO?” is like asking “payment workflow or database transaction?” One describes an overall process; the other describes a mechanism that may be used within it.

Error 3: Comparing Algorithm Results Without Drawing the System Topology ​

If the comparison does not account for Reward Model, Critic, Rollout, environment, and communication costs, the algorithmic gain is not the complete engineering gain.

Error 4: Copying Someone Else's Answer Without Copying Their Constraints ​

A team sees GLM-5.2 use PPO and follows it, even though it has no long, variable-length, compacted trajectories and needs no Token-level credit assignment. It has copied the conclusion, not the reasoning.

Error 5: Declaring the Architecture Valid Because the Score Rose ​

A change may improve the target benchmark while harming general capability, safety, training stability, or efficiency per unit of compute. Model promotion requires a body of evidence, not one flattering number.


🎯 Quick Check ​

🤔In a reinforcement-learning training system, what is the main difference between a Reward Model and a Critic?
🤔A team is debating whether RLHF or PPO is better. What should it do first?
🤔A math task can cheaply generate multiple answers for the same problem, verify them automatically, and usually produces clear within-group Reward differences. The team also wants to avoid Critic cost. Which route is a more natural match by assumption?
🤔Why is the GLM-5.2 case an architectural decision rather than merely an algorithm rename?

Chapter Summary ​

  • Policy is the model strategy being optimized; Actor is its role when executing that strategy in a Rollout.
  • Rollout is the actual attempt; Trajectory is the complete experience record it leaves behind.
  • Reward evaluates results; a Critic estimates future Value; Advantage measures how actual performance compares with a baseline.
  • SFT learns demonstrations, DPO learns offline preferences, GRPO uses same-prompt group comparison to eliminate the Critic, and Critic-based PPO pays for an additional value model to gain single-trajectory learning and fine-grained attribution.
  • RLHF, RLAIF, and verifiable rewards mainly describe feedback sources or training processes; they are not simply mutually exclusive with PPO or GRPO.
  • Training methods are not an algorithm leaderboard. The correct sequence is goal → material → signal → attribution → optimization → components → system → evidence.
  • The lesson of GLM-5.2 is not that PPO beats GRPO. Long, variable-length, compacted trajectories no longer fit regular group comparison naturally, so the team accepted Critic cost in exchange for single-trajectory learning and Token-level attribution.
  • Reuse the reasoning; do not copy the answer. Every training decision must state candidate assumptions, system impact, ablation evidence, and exit conditions.

Question for the next chapter: once the decision coordinate system is established, the first fundamental pre-training decision is not which framework to choose, but how to divide a finite compute budget among model parameters, training Tokens, and higher-quality data.


Basic Glossary ​

TermA beginner's first intuition
PolicyThe current model's probability rules for “what to do next”
ActorThe role of the Policy when it actually generates answers or performs actions
RolloutLetting the current Policy attempt a complete task once
TrajectoryThe input, output, actions, and environment results recorded from an attempt
RewardA score assigned to a result that has already been produced
VerifierA component that calculates Reward using rules, tests, or environment state
Reward ModelA model that learns to score results from preference data
ValueThe expected final return from continuing at the current state
CriticA model that learns and estimates Value
AdvantageHow much better or worse actual performance was than an expected baseline
Credit AssignmentAttributing the final result to intermediate Turns, Actions, or Tokens
Reference PolicyA frozen reference model that constrains the current Policy from drifting too far
On-policyTraining data comes from the current Policy or a very recent version
Policy LagThe Policy version generating data lags behind the version being trained
CompactionCompressing long-horizon history, potentially splitting a very long trajectory into sub-trajectories
Fine-tuningContinuing training on an existing model instead of starting from scratch
Offline TrainingUsing fixed data prepared before training begins
Online RolloutContinuously generating new experience with the current Policy during training
CheckpointModel parameters and recovery state saved at a point during training
Reward HackingExploiting scoring-rule loopholes to obtain a high score without solving the real task
Ablation StudyChanging one factor at a time to identify where an improvement came from

💬 Comments