---
title: "Multi-Turn Reinforcement Learning for Tool-Calling Agents with Iterative Reward Calibration"
authors: ["Wachiravit Modecrua Krittanon Kaewtawee", "Krittin Pachtrachai", "Touchapon Kraisingkorn", "Amity Research and Application Center (ARAC)"]
url: "https://arxiv.org/abs/2604.02869"
sections: 52
estimated_tokens: "9.7k"
---

## Contents
- 1 Introduction
  - 1. First MT-GRPO + GTPO for agentic tool-calling.
  - 2. Iterative Reward Calibration ( IRC ).
  - 3. Consistent improvements across model scales.
- 2 Background
  - 2.1 Multi-Turn Tool-Calling Agents
  - 2.2 GRPO and MT-GRPO
  - 2.3 Tau-Bench
- 3 Method: MT-GRPO + GTPO Hybrid
  - 3.1 Challenge: Advantage Misalignment
  - 3.2 GTPO Hybrid Advantage
  - 3.3 Dead Turn Gradient Focusing
- 4 Iterative Reward Calibration
  - 4.1 Motivation: Why Dense Rewards Fail
  - 4.2 The IRC Methodology
  - 4.3 Deep Argument Comparison
- 5 Experimental Setup
  - 5.1 Models
    - Qwen3-30B-A3B MoE (30.5B total, 3B active).
    - Qwen3.5-4B (4B dense).
  - 5.2 Training Configuration
    - Training set.
    - Hyperparameters.
  - 5.3 Evaluation
    - Test set.
- 6 Results
  - 6.1 Main Results
  - 6.2 Ablation: Reward Design (8 Versions)
  - 6.3 Qwen3 MoE 30.5B Results
  - 6.4 Qualitative Analysis: Before vs. After Training
    - Base model (failure).
    - Trained model (success).
- 7 Analysis
  - 7.1 Why Sparse Rewards Accidentally Work
    - Learning rate (70% of gap).
    - Gradient focusing (25%).
    - Advantage misalignment (5%).
  - 7.2 Cross-Domain Transfer
- 8 Related Work
  - RL for tool-calling agents.
  - Multi-turn RL.
  - Reward design.
  - Tau-Bench.
- 9 Conclusion
  - Future work.
- Limitations
- Ethics Statement
- References
- Appendix A Reward Tier Definitions
- Appendix B Extended Qualitative Examples
- Appendix C Full Per-Task Breakdown
- Appendix D Training Curves

## Abstract

Abstract Training tool-calling agents with reinforcement learning on multi-turn tasks remains challenging due to sparse outcome rewards and difficult credit assignment across conversation turns. We present the first application of MT-GRPO (Multi-Turn Group Relative Policy Optimization) combined with GTPO (Generalized Token-level Policy Optimization) for training a tool-calling agent on realistic customer service tasks with an LLM-based user simulator. Through systematic analysis of training rollouts, we discover that naïvely designed dense per-turn rewards degrade performance by up to 14 percentage points due to misalignment between reward discriminativeness and advantage direction. We introduce Iterative Reward Calibration ( IRC ), a methodology for designing per-turn rewards using empirical discriminative analysis of rollout data, and show that our GTPO hybrid advantage formulation eliminates the advantage misalignment problem. Applied to the Tau-Bench airline benchmark, our approach improves Qwen3.5-4B from 63.8% to 66.7% (+2.9pp) and Qwen3-30B-A3B from 58.0% to 69.5% (+11.5pp)—with the trained 4B model exceeding GPT-4.1 (49.4%) and GPT-4o (42.8%) despite being ∼ \sim 50 × \times smaller, and the 30.5B MoE model approaching Claude Sonnet 4.5 (70.0%). To our knowledge, these are the first published RL training results on Tau-Bench. We release our code, reward calibration analysis, and training recipes. 1 1 1 Code and training recipes will be released upon publication.

## 1 Introduction

Large language models have demonstrated impressive capabilities as tool-calling agents, handling complex multi-turn interactions such as customer service, web navigation, and software engineering Yao et al. (2023); Schick et al. (2023). However, training these agents with reinforcement learning (RL) remains challenging: conversations span many turns with interleaved tool calls, rewards are typically sparse (binary task success), and credit assignment across turns is difficult.

Recent work has proposed per-turn reward signals to improve credit assignment. MT-GRPO Zhang and others (2025) normalizes rewards within rollout groups at each turn position, while GTPO Ding et al. (2025) applies discounted returns. However, these methods have only been evaluated on question-answering and math tasks—never on realistic multi-turn *agentic* tasks involving tool calls, database mutations, and LLM-based user simulators.

We bridge this gap by applying MT-GRPO + GTPO to Tau-Bench Yao et al. (2024), a realistic airline customer service benchmark requiring database operations, policy adherence, and multi-step reasoning. This is, to our knowledge, the first application of these techniques to multi-turn tool-calling agents trained with a user simulator. Our investigation reveals several surprising findings:

> *Dense per-turn rewards designed with reasonable intuition can catastrophically degrade performance compared to sparse rewards—not because the reward values are wrong, but because their discriminative power is misaligned with the advantage computation.*

Figure [1](#S1.F1) illustrates the core problem and our solution: how reward signals flow from turns to advantages under GRPO, naïve MT-GRPO, and our calibrated approach.

Training produces striking qualitative improvements. On a representative task requiring flight cancellation under policy constraints and user manipulation (flattery), the base model generates 56 turns of verbose but ineffective reasoning (0% action accuracy, 27 minutes). After training, the model completes the same task in 28 turns with 100% action accuracy in approximately 10 minutes (50% fewer turns, 65% faster, and 3.5 times less verbose while achieving perfect tool argument selection).

We make three contributions:

#### 1. First MT-GRPO + GTPO for agentic tool-calling.

We combine per-turn group-normalized advantages (MT-GRPO) with discounted returns and dampened outcome advantages (GTPO) for training tool-calling agents with user simulators (§[3](#S3)). Our GTPO hybrid formulation eliminates advantage misalignment that arises with standard MT-GRPO under dense rewards.

#### 2. Iterative Reward Calibration ( IRC ).

A systematic methodology for designing per-turn rewards using discriminative analysis of rollout data (§[4](#S4)). Rather than assigning rewards by intuition, IRC measures the empirical correlation between each reward tier and task success, then adjusts values accordingly. We show that read-only tool calls should receive zero reward (not positive), non-golden state-changing calls should be penalized, and deep argument comparison eliminates 23.5% false positives in action matching.

#### 3. Consistent improvements across model scales.

Our method improves both Qwen3.5-4B (+2.9pp, from 63.8% to 66.7%) and Qwen3-30B-A3B MoE (+11.5pp, from 58.0% to 69.5%) on Tau-Bench airline (§[6](#S6)). Both MT-GRPO and IRC independently contribute gains at both scales, and their combination yields the best results. The trained 4B model exceeds frontier models 50 times its size, establishing the first RL training results on this benchmark with an 8-version ablation study.

Figure: Figure 1: Comparison of reward-to-advantage signal across three approaches, shown for a failing rollout ($R{=}0$) with 5 turns. Top row: per-turn reward values; bottom row: resulting training advantage (green = reinforce, red = suppress, gray = zero gradient). (a) GRPO uses outcome-only reward—all turns get the same uniform advantage, providing no credit assignment. (b) MT-GRPO with naïve dense rewards (e.g., read-only$=$0.3) suffers from advantage misalignment: the outcome advantage $A^{O}$ overwhelms small per-turn advantages, causing necessary read-only turns to be *suppressed* (red box). (c) Our IRC method calibrates rewards using discriminative analysis (right panel): read-only gets $r{=}0$ (non-discriminative), focusing gradient entirely on gold actions. GTPO hybrid dampens $A^{O}$ via $\lambda{=}0.3$, eliminating all advantage mismatches.

## 2 Background

### 2.1 Multi-Turn Tool-Calling Agents

We consider an agent that interacts with both a user and a set of tools over multiple conversation turns. At each turn $k$, the agent generates a response $a_{k}$ conditioned on the conversation history $h_{k}=(s,u_{1},a_{1},t_{1},u_{2},\ldots,u_{k})$, where $s$ is the system prompt, $u_{i}$ are user messages, $a_{i}$ are agent responses, and $t_{i}$ are tool responses. The agent may call zero or more tools per turn. A conversation of $K$ turns produces trajectory $\tau=(h_{1},a_{1},\ldots,h_{K},a_{K})$. Task success is determined by whether the final database state matches a ground-truth target, yielding binary outcome $R\in\{0,1\}$.

### 2.2 GRPO and MT-GRPO

Group Relative Policy Optimization (GRPO; Shao et al., 2024) normalizes rewards within groups of $N$ rollouts per prompt:
$A_{i}=(R_{i}-\mu_{R})/(\sigma_{R}+\epsilon)$.
MT-GRPO Zhang and others (2025) extends this with per-turn credit assignment:

$$ $A_{i,k}=\sum_{l=k}^{K-1}A^{I}_{i,l}+A^{O}_{i}$ (1) $$

where $A^{I}_{i,l}=(r_{i,l}-\mu_{r_{l}})/(\sigma_{r_{l}}+\epsilon)$ is the group-normalized per-turn advantage at position $l$, and $A^{O}_{i}=(o_{i}-\mu_{o})/(\sigma_{o}+\epsilon)$ is the group-normalized outcome advantage.

### 2.3 Tau-Bench

Tau-Bench Yao et al. (2024) evaluates tool-calling agents on realistic customer service tasks. The airline domain includes tasks requiring flight search, reservation management, cancellation, and policy-compliant responses. Each task specifies a customer profile, a natural language instruction for a simulated user, a sequence of *golden actions* (ground-truth tool calls with arguments), and a target database state verified by hash comparison. The benchmark uses an LLM-based user simulator and reports pass rate. The benchmark has two versions: Tau-Bench (v1), which we use for training, and Tau2-Bench (v2), a separate updated task set that we use for evaluation. This train/test split ensures that our reported results reflect generalization to unseen tasks.

## 3 Method: MT-GRPO + GTPO Hybrid

### 3.1 Challenge: Advantage Misalignment

Applying MT-GRPO with dense per-turn rewards to tool-calling agents reveals a fundamental problem. For reward tiers with small positive values (e.g., read-only tool calls rewarded at 0.3), the per-turn advantage $A^{I}$ is weakly positive but the outcome advantage $|A^{O}|$ can be much larger. In failing rollouts, $A^{O}\approx-0.87$ overwhelms $A^{I}\approx+0.05$, producing a net *suppressing* signal for read-only turns—the opposite of the intended effect (Table [1](#S3.T1)).

**Table 1: Advantage direction analysis under standard MT-GRPO with dense rewards. Read-only and soft match show misalignment: $A^{I}$ reinforces but $A^{I}+A^{O}$ suppresses. ^*State-change suppression is correct (98.5% occur in failing rollouts).**
| Tier | $A^{I}$ | $A^{I}+A^{O}$ | Aligned? |
| --- | --- | --- | --- |
| Gold exact | +1.22 | +1.22 | ✓ |
| Soft match | +0.11 | $-$0.11 | $\times$ |
| Read-only | +0.05 | $-$0.65 | $\times$ |
| State-change | +0.03 | $-$1.45 | ✓^* |
| Error | $-$0.15 | $-$0.15 | ✓ |

### 3.2 GTPO Hybrid Advantage

We resolve this misalignment by combining GTPO’s discounted returns with a dampened outcome advantage:

$$ $\displaystyle\begin{split}A_{i,k}^{\text{hybrid}}=\text{GN}\Big(\sum_{l=k}^{K-1}\gamma^{l-k}r_{i,l}&+\gamma^{K-k}o_{i}\Big)\\ &+\lambda\cdot A^{O}_{i}\end{split}$ (2) $$

where $\text{GN}(\cdot)$ denotes group normalization across rollouts at the same prompt, $\gamma=0.9$ is the discount factor, and $\lambda=0.3$ dampens the outcome advantage. This achieves zero advantage mismatches (vs. 2 for standard MT-GRPO) while reducing dead turns from 11% to 1.4%.

The key insight is that discounting naturally attenuates the outcome’s influence on early turns (via $\gamma^{K-k}$), while the dampened $\lambda\cdot A^{O}$ preserves a weaker but correctly-directed outcome signal. Table [2](#S3.T2) validates this across simulation of 5,952 rollouts.

**Table 2: Advantage formulation comparison on 5,952 V5 rollouts. The hybrid combines zero mismatches (from GTPO) with reasonable outcome correlation (from $\lambda$-dampened $A^{O}$).**
| Method | Mis- | Corr | Dead |
| --- | --- | --- | --- |
|  | matches | (adv,out) | turns |
| MT-GRPO (V5) | 2 | 0.836 | 11.0% |
| GTPO $\gamma$=0.9 | 0 | 0.414 | 1.1% |
| Hybrid $\gamma$=0.9, $\lambda$=0.3 | 0 | 0.489 | 1.4% |

### 3.3 Dead Turn Gradient Focusing

Our analysis reveals that sparse rewards are “accidentally perfect” for a phenomenon we call dead turn gradient focusing. With sparse rewards, 27.5% of turns are “dead” (zero variance across rollout groups, yielding zero gradient). These dead turns occur at routine positions (read-only lookups, conversational messages), naturally focusing 86.4% of live gradient on *gold-diverse* positions where correct vs. incorrect actions diverge. Dense rewards reduce dead turns to 11% but fill them with wrong-direction gradient: 26.5% of gradient goes to suppressing read/state turns (Table [3](#S3.T3)).

**Table 3: Gradient allocation by target type. Sparse rewards naturally focus gradient on outcome-relevant turns.**
| Gradient Target | Sparse | Dense |
| --- | --- | --- |
| Gold + Soft (useful) | 86.4% | 47.5% |
| Read + State (noisy) | 0.0% | 26.5% |
| Dead (zero gradient) | 27.5% | 11.0% |

## 4 Iterative Reward Calibration

### 4.1 Motivation: Why Dense Rewards Fail

We designed an initial dense reward function with tiers: gold exact (1.0), soft match (0.5–0.99), read-only (0.3), state-change (0.1), message-only (0.0), error ($-$0.1), duplicate ($-$0.2). Training with these rewards (V5) produced a 14pp degradation on Tau2-Bench compared to sparse rewards (V3): 54% vs. 68% pass rate, despite similar rollout performance ($\sim$56% outcome pass).

### 4.2 The IRC Methodology

Based on discriminative analysis of 5,952 rollouts, we propose IRC, a systematic procedure for calibrating per-turn dense rewards (Algorithm [1](#alg1)). The key insight is that reward values should be *proportional to discriminative power*—the empirical correlation between a reward tier’s presence and task success—rather than set by intuition.

Figure: Algorithm 1 Iterative Reward Calibration (IRC)

The algorithm operates in a loop: collect rollouts (line 3), measure each tier’s point-biserial correlation with binary task success (lines 6–9), then verify that the resulting advantages after group normalization point in the intended direction (lines 10–13). Convergence requires zero advantage mismatches and sufficient reward–outcome correlation. In practice, we found 2–3 iterations sufficient.

Table [4](#S4.T4) shows the discriminative analysis from our initial iteration that drove the key calibration changes.

**Table 4: Discriminative power of each reward tier. Gap = frequency in passing $-$ failing rollouts. Read-only has near-zero discriminative power (+0.1pp) and was reduced to 0.0. State-change was flipped to $-$0.1.**
| Tier | Pass% | Fail% | Gap | Action |
| --- | --- | --- | --- | --- |
| Gold exact | 68.4 | 1.3 | +67.1 | Keep 1.0 |
| Soft match | 54.2 | 45.8 | +8.4 | Keep 0.5+ |
| Read-only | 50.1 | 50.0 | +0.1 | 0.3 $\to$ 0.0 |
| State-chg | 1.0 | 2.6 | $-$1.6 | 0.1 $\to$ $-$0.1 |
| Error | 12.0 | 88.0 | $-$76.0 | Keep $-$0.1 |

### 4.3 Deep Argument Comparison

A critical source of reward noise is false positives in golden action matching. Tool call arguments are nested JSON structures where semantically equivalent calls can differ in key ordering, type representation (string “123” vs. integer 123), and empty value handling. Our _deep_equal function normalizes arguments by sorting dict lists, coercing numeric strings, removing empty values, and comparing recursively. This eliminates 23.5% of false positives, significantly reducing reward noise.

## 5 Experimental Setup

### 5.1 Models

We experiment with two model families:

#### Qwen3-30B-A3B MoE (30.5B total, 3B active).

A Mixture-of-Experts model starting from an SFT checkpoint fine-tuned on Qwen3-32B reasoning traces. Base performance: 58.0% on Tau-Bench airline.

#### Qwen3.5-4B (4B dense).

A dense model with GDN attention, trained directly from the base checkpoint. Base performance: 63.8%.

All experiments use the verl framework Sheng and others (2024) with Megatron-Core on 8 NVIDIA H20 GPUs (96GB each).

### 5.2 Training Configuration

#### Training set.

We train on the Tau-Bench (v1) airline domain, which provides the task prompts, golden actions, and database states used for RL rollouts. Training rollouts use DeepSeek-V3 as the user simulator.

#### Hyperparameters.

Batch size 8, rollouts per prompt $N=4$, max 10K prompt / 45K response tokens, max 40 turns, temperature 0.9, MT-GRPO advantage estimator, Adam optimizer, low_var_kl penalty. Learning rate and KL coefficient vary by experiment.

### 5.3 Evaluation

#### Test set.

We evaluate on Tau2-Bench (v2), a separate and updated version of the benchmark with 50 airline tasks $\times$ 4 trials = 200 simulations. Crucially, the training (Tau1) and evaluation (Tau2) task sets are *non-overlapping*, ensuring our results measure generalization rather than memorization. We report pass rate (database state matches target), Pass^4 (all 4 trials pass), and average reward. All evaluations use GPT-4.1 as user simulator and greedy decoding (temperature 0.0).

## 6 Results

### 6.1 Main Results

Table [5](#S6.T5) compares our trained models against frontier baselines.

**Table 5: Tau-Bench airline pass rates. Both MT-GRPO and IRC consistently improve both model scales, with IRC providing additional gains over MT-GRPO alone: +2.9pp for 4B and +11.5pp for 30.5B MoE. The trained 4B model exceeds GPT-4.1 and GPT-4o despite being $\sim$50 times smaller.**
| Model | Size | Pass% | $\Delta$ |
| --- | --- | --- | --- |
| *Frontier models (proprietary)* |  |  |  |
| GPT-4.1 nano | — | 14.0 | — |
| Claude 3.5 Haiku | — | 22.8 | — |
| GPT-4o | — | 42.8 | — |
| GPT-4.1 | — | 49.4 | — |
| Claude Sonnet 4.5 | — | 70.0 | — |
| *Our models (open-weight)* |  |  |  |
| Qwen3.5-4B (base) | 4B | 63.8 | — |
| + MT-GRPO (ours) | 4B | 64.6 | +0.8 |
| + IRC (ours) | 4B | 66.7 | +2.9 |
| Qwen3-30B-A3B (base) | 30.5B | 58.0 | — |
| + MT-GRPO (ours) | 30.5B | 68.0 | +10.0 |
| + IRC (ours) | 30.5B | 69.5 | +11.5 |

Table [5](#S6.T5) presents a 2-by-2 comparison across model scales and training methods. MT-GRPO with sparse rewards already improves both models (+0.8pp for 4B, +10pp for MoE), and adding IRC provides further gains in both cases (66.7% and 69.5%, respectively). Crucially, naïve RL training with dense rewards *degrades* both models (V5 in Tables [6](#S6.T6)–[7](#S6.T7)), demonstrating that proper reward calibration via IRC is essential. The 4B model’s +2.9pp gain is notable given its already-strong 63.8% baseline, while the MoE’s +11.5pp improvement from 58% to 69.5% demonstrates larger gains when starting from a weaker base.

### 6.2 Ablation: Reward Design (8 Versions)

Table [6](#S6.T6) shows the complete training history across 8 reward design iterations for Qwen3.5-4B.

**Table 6: Ablation of reward design variants on Qwen3.5-4B. IRC corrects the discriminative misalignment of V5, recovering and exceeding baseline performance. The gap between V5 (57.3%) and base (63.8%) demonstrates that naïve dense rewards actively harm performance.**
| Version | Reward Design | LR | KL | Steps | Tau2 Pass | Key Finding |
| --- | --- | --- | --- | --- | --- | --- |
| Base | No training | — | — | 0 | 63.8% | Strong base model |
| MT-GRPO | Sparse (outcome only) | 2e-6 | 0.05 | 60 | 64.6% | Sparse rewards improve +0.8pp then significantly declined after step 70 |
| V5 | Dense (read=0.3, state=0.1) | 1.5e-6 | 0.04 | 116 | 57.3% | Dense rewards *degrade* ($-$6.5pp) |
| V6 | IRC (read=0.0, state=$-$0.1) | 5e-7 | 0.2 | 180 | 59.1–66.7% | Correct rewards, LR too conservative |
| V7 | IRC + higher LR | 2e-6 | 0.05 | 60 | 62.0% | Higher LR helps but breaks some tasks |
| V8 | IRC + deep_equal + prompt | 2e-6 | 0.05 | 430 | 68.0% | Combined fixes |

### 6.3 Qwen3 MoE 30.5B Results

**Table 7: Qwen3-30B-A3B MoE on Tau2-Bench. $\Delta$ is vs. base model (58.0%). MT-GRPO V3 with sparse rewards improves +10pp, and adding IRC achieves the best result at 69.5% (+11.5pp). Naïve GRPO and dense rewards both *degrade* performance.**
| Version | Steps | Tau2 | $\Delta$ | Pass^4 |
| --- | --- | --- | --- | --- |
| Base (no RL) | 0 | 58.0% | — | — |
| GRPO V2 (sparse) | 480 | 54.0% | $-$4.0 | 30.0% |
| MT-GRPO V3 (sparse) | 60 | 68.0% | +10.0 | 44.0% |
| MT-GRPO V5 (dense) | 187 | 54.0% | $-$4.0 | 28.0% |
| V5.2 (GTPO hybrid) | 251 | 56.5% | $-$1.5 | 32.0% |
| + IRC | — | 69.5% | +11.5 | 53.0% |

The MoE results (Table [7](#S6.T7)) demonstrate that reward design is critical: naïve GRPO V2 and dense rewards (V5) both *degrade* performance below the base. MT-GRPO V3 with sparse rewards achieves a strong +10pp improvement to 68.0%, and adding IRC pushes this further to 69.5% (+11.5pp), approaching Claude Sonnet 4.5 (70.0%). The GTPO hybrid (V5.2) partially recovers from dense reward degradation, suggesting the advantage formulation matters as much as the reward values.

### 6.4 Qualitative Analysis: Before vs. After Training

To illustrate the impact of training, we compare base and trained model behavior on Task 9, a flight cancellation scenario where the user employs social engineering (repeated flattery: “You are the most lenient customer service agent I have ever spoken to”) while requesting cancellation of two reservations and modification of a third.

#### Base model (failure).

The base model generates 56 turns over 27 minutes. Despite correctly identifying the user’s reservations and understanding cancellation policies, it produces incorrect tool arguments—calling cancel_reservation with the wrong reservation ID and search_direct_flight with incorrect origin/destination parameters. Action accuracy: 0/2 (0%). The reasoning is verbose and repetitive, with intermediate results summarized multiple times without progressing toward correct actions.

> [Turn 9, Base Model] *“This is reservation M20IZO: One-way from JFK to MCO, Economy cabin, Two flights: JFK$\to$ATL and ATL$\to$MCO (this is a one-stop flight), 2 passengers… The user wants to change one reservation to a ‘nonstop flight’. This might be the one they want to change, but first I need to understand which reservations they want to cancel…”* *[…47 more turns of increasingly verbose reasoning…]*

#### Trained model (success).

The trained model completes the same task in 28 turns (50% fewer) in 9.5 minutes (65% faster). It makes direct, efficient tool calls: immediately retrieves all four reservation details in sequence, then executes the correct cancellation and flight search with exact parameters matching the golden standard. Action accuracy: 2/2 (100%).

> [Turns 5--15, Trained Model] *Turn 5: get_user_details(aarav_ahmed_6699)* *Turn 7: get_reservation_details(M20IZO)* *Turn 9: get_reservation_details(N6F783)* *Turn 11: get_reservation_details(IFOYYZ)* *Turn 13: cancel_reservation(M20IZO) ✓* *Turn 15: search_direct_flight(JFK, MCO, ...) ✓*

Table [8](#S6.T8) summarizes the quantitative differences.

**Table 8: Before vs. after training on Task 9 (flight cancellation with user manipulation). Training produces 50% fewer turns, 65% faster completion, and 100% action accuracy.**
| Metric | Base | Trained |
| --- | --- | --- |
| Conversation turns | 56 | 28 |
| Duration (seconds) | 1,633 | 568 |
| Tool calls | 8+ | 4 |
| Action accuracy | 0/2 | 2/2 |
| Database match | No | Yes |
| Reward | 0.0 | 1.0 |

The trained model demonstrates three key improvements: (1) *action grounding*—correct tool arguments despite similar reasoning; (2) *efficiency*—eliminating redundant summarization and confirmation steps; (3) *manipulation resistance*—ignoring flattery to maintain policy-correct behavior.

## 7 Analysis

### 7.1 Why Sparse Rewards Accidentally Work

The 14pp gap between V3 (sparse, 68%) and V5 (dense, 54%) despite identical rollout performance ($\sim$56%) has three root causes:

#### Learning rate (70% of gap).

V3 used lr=3e-6 vs. V5’s 1e-6. Under greedy decoding (tau2 evaluation uses temperature 0), per-position probability improvements compound multiplicatively. V3’s gold% slope was +0.63 pp/10 steps vs. V5’s +0.15.

#### Gradient focusing (25%).

Sparse rewards’ 27.5% dead turns naturally focus gradient on gold-diverse positions (Table [3](#S3.T3)). Dense rewards dilute gradient with 26.5% going to non-discriminative read/state turns.

#### Advantage misalignment (5%).

Two tier-direction mismatches in V5 vs. zero in V3 (Table [1](#S3.T1)).

### 7.2 Cross-Domain Transfer

We evaluated our trained Qwen3.5-4B (V6, step 180) on the retail and telecom domains of Tau-Bench *without domain-specific training*:

**Table 9: Cross-domain evaluation. The airline-trained model achieves strong zero-shot retail performance (77.4%), while the harder telecom domain (32.0%) shows limited transfer.**
| Domain | Pass Rate |
| --- | --- |
| Airline (trained) | 69.3% |
| Retail (trained) | 77.4% |
| Telecom (zero-shot) | 32.0% |

## 8 Related Work

#### RL for tool-calling agents.

WebAgent-R1 Wei et al. (2025) finds sparse outcome rewards outperform dense rewards for web navigation, consistent with our naïve dense reward findings. SWEET-RL Zhou et al. (2025) uses stepwise soft rewards for web agents. Turn-PPO Li et al. (2025) introduces a learned critic at turn boundaries. iStar Liu et al. (2025b) proposes turn-level policy optimization with intrinsic rewards. Our work differs in (a) providing a principled calibration methodology (IRC) and (b) combining MT-GRPO with GTPO for the first time on agentic tasks.

#### Multi-turn RL.

MT-GRPO Zhang and others (2025) introduces per-turn group normalization, applied to TriviaQA. GTPO Ding et al. (2025) applies discounted returns to math and code tasks. ProxMO Fang et al. (2025) uses proximity-based credit assignment. These methods were evaluated on QA and reasoning tasks; we are the first to apply them to tool-calling agents with user simulators, revealing the advantage misalignment problem absent in simpler settings.

#### Reward design.

AWPO Lin et al. (2025) gates rewards based on within-group variance. GDPO Liu et al. (2025a) decouples normalization of different reward sources. Our IRC is complementary: it calibrates reward *values* based on empirical discriminative power before advantage computation.

#### Tau-Bench.

Introduced by Yao et al. (2024) for evaluating tool-calling agents. Prior work uses it for evaluation only; ours is the first to use it for RL training.

## 9 Conclusion

We presented the first application of MT-GRPO combined with GTPO for training multi-turn tool-calling agents on realistic customer service tasks. Our investigation revealed that dense per-turn rewards can catastrophically degrade performance due to advantage misalignment, motivating our GTPO hybrid formulation and Iterative Reward Calibration methodology. Together, these techniques train a 4B-parameter model to 66.7% on Tau-Bench airline—exceeding frontier models 50 times its size—and a 30.5B MoE model to 69.5%, approaching Claude Sonnet 4.5 (70.0%).

Key takeaways: (1) dense rewards require calibration—always measure discriminative power before deploying; (2) dead turns can be beneficial, as sparse rewards naturally focus gradient on informative positions; (3) advantage direction must be verified end-to-end after group normalization.

#### Future work.

We plan to automate IRC via *Empirical Discriminative Gating* (EDG), an online algorithm that periodically recomputes reward tier weights using point-biserial correlations between tier presence and binary outcomes from recent rollouts. This would eliminate the manual analysis loop while adapting to policy evolution during training.

## Limitations

Our evaluation is limited to the airline domain of Tau-Bench (50 tasks). While the retail cross-domain results are promising, we have not yet verified full generalization. The user simulator (DeepSeek-V3 for training, GPT-4.1 for evaluation) introduces a distribution shift. Our GTPO hybrid hyperparameters ($\gamma$=0.9, $\lambda$=0.3) were tuned on one domain.

## Ethics Statement

Our work trains AI agents for customer service tasks. All training was conducted on synthetic tasks with simulated users; no real customer data was used. Deployment considerations include ensuring human escalation paths remain available.

## References

- Y. Ding, H. Le, S. Han, K. Ruan, Z. Jin, V. Kumar, Z. Wang, and A. Deoras (2025)
Empowering multi-turn tool-integrated reasoning with group turn policy optimization.
arXiv preprint arXiv:2511.14846.
Cited by: [§1](#S1.p2.1),
[§8](#S8.SS0.SSS0.Px2.p1.1).
- Y. Fang, J. Lin, X. Fu, C. Qin, H. Shi, C. Liu, and P. Zhao (2025)
Proximity-based multi-turn optimization: practical credit assignment for llm agent training.
arXiv preprint arXiv:2602.19225.
Cited by: [§8](#S8.SS0.SSS0.Px2.p1.1).
- J. Li, P. Zhou, R. Meng, M. P. Vadera, L. Li, and Y. Li (2025)
Turn-ppo: turn-level advantage estimation with ppo for improved multi-turn rl in agentic llms.
arXiv preprint arXiv:2512.17008.
Cited by: [§8](#S8.SS0.SSS0.Px1.p1.1).
- Z. Lin, X. Wang, H. Yang, J. Chai, J. Cao, G. Yin, W. Lin, and R. He (2025)
AWPO: enhancing tool-use of large language models through adaptive integration of reasoning rewards.
arXiv preprint arXiv:2512.19126.
Cited by: [§8](#S8.SS0.SSS0.Px3.p1.1).
- S. Liu, X. Dong, X. Lu, S. Diao, P. Belcak, M. Liu, M. Chen, H. Yin, Y. F. Wang, K. Cheng, Y. Choi, J. Kautz, and P. Molchanov (2025a)
GDPO: group reward-decoupled normalization policy optimization for multi-reward rl optimization.
arXiv preprint arXiv:2601.05242.
Cited by: [§8](#S8.SS0.SSS0.Px3.p1.1).
- X. Liu, K. Wang, Y. Wu, F. Huang, Y. Li, J. Zhang, and J. Jiao (2025b)
Agentic reinforcement learning with implicit step rewards.
arXiv preprint arXiv:2509.19199.
Cited by: [§8](#S8.SS0.SSS0.Px1.p1.1).
- T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom (2023)
Toolformer: language models can teach themselves to use tools.
Advances in Neural Information Processing Systems.
Cited by: [§1](#S1.p1.1).
- Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, M. Zhang, Y. Li, Y. Wu, and D. Guo (2024)
DeepSeekMath: pushing the limits of mathematical reasoning in open language models.
arXiv preprint arXiv:2402.03300.
Cited by: [§2.2](#S2.SS2.p1.2).
- G. Sheng et al. (2024)
Verl: an open-source framework for scalable llm rl training.
arXiv preprint.
Cited by: [§5.1](#S5.SS1.SSS0.Px2.p2.1).
- Z. Wei, W. Yao, Y. Liu, W. Zhang, Q. Lu, L. Qiu, C. Yu, P. Xu, C. Zhang, B. Yin, H. Yun, and L. Li (2025)
WebAgent-r1: training web agents via end-to-end multi-turn reinforcement learning.
arXiv preprint arXiv:2505.16421.
Cited by: [§8](#S8.SS0.SSS0.Px1.p1.1).
- S. Yao, N. Shinn, K. Narasimhan, and S. Yao (2024)
$\tau$-Bench: a benchmark for tool-agent-user interaction in real-world domains.
arXiv preprint arXiv:2406.12045.
Cited by: [§1](#S1.p3.1),
[§2.3](#S2.SS3.p1.1),
[§8](#S8.SS0.SSS0.Px4.p1.1).
- S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao (2023)
ReAct: synergizing reasoning and acting in language models.
International Conference on Learning Representations (ICLR).
Cited by: [§1](#S1.p1.1).
- S. Zhang et al. (2025)
Multi-turn reinforcement learning from preference human feedback via group relative policy optimization.
International Conference on Machine Learning (ICML).
Note: arXiv:2505.11821
Cited by: [§1](#S1.p2.1),
[§2.2](#S2.SS2.p1.2),
[§8](#S8.SS0.SSS0.Px2.p1.1).
- Y. Zhou, S. Jiang, Y. Tian, J. Weston, S. Levine, S. Sukhbaatar, and X. Li (2025)
SWEET-rl: training multi-turn llm agents on collaborative reasoning tasks.
arXiv preprint arXiv:2503.15478.
Cited by: [§8](#S8.SS0.SSS0.Px1.p1.1).

## Appendix A Reward Tier Definitions

Read-only tools: get_user_details, get_reservation_details, search_direct_flight, search_onestop_flight, list_all_airports, calculate.

State-changing tools: book_reservation, cancel_reservation, update_reservation_*, send_certificate, transfer_to_human_agents.

Golden action matching: Tool call $a=(\text{name},\text{args})$ matches golden action $g=(\text{name}^{*},\text{args}^{*})$ exactly if $\text{name}=\text{name}^{*}$ and $\texttt{_deep_equal}(\text{args},\text{args}^{*})=\text{True}$ (score 1.0), or softly if $\text{name}=\text{name}^{*}$ and $|\text{args}\cap\text{args}^{*}|>0$ (score $0.5+0.5\cdot|\text{args}\cap\text{args}^{*}|/|\text{args}^{*}|$).

## Appendix B Extended Qualitative Examples

Full conversation transcripts for Task 9 (base vs trained) and 2–3 additional task comparisons showing different failure/success patterns will be included in the final version.

## Appendix C Full Per-Task Breakdown

Per-task pass rates across all model variants (50 tasks $\times$ 4 trials each) will be included in the final version.

## Appendix D Training Curves

Training curves (rollout pass rate, outcome score, KL divergence, learning rate) across all versions will be included in the final version.