Title: 2604.08065
ArXiv: 2604.08065

Multimodal Latent Reasoning via Predictive Embeddings

Title:

Content selection saved. Describe the issue below:

Description:

License: CC BY 4.0

arXiv:2604.08065v1 [cs.LG] 09 Apr 2026

Multimodal Latent Reasoning via Predictive Embeddings

Ashutosh Adhikari & Mirella Lapata

School of Informatics

University of Edimburgh

{ashutosh.adhikari, mlap}@ed.ac.uk

Abstract

Tool-augmented multimodal reasoning enables visual language models (VLMs) to improve perception by interacting with external tools (e.g., cropping, depth estimation). However, such approaches incur substantial inference overhead, require specialized supervision, and are prone to erroneous tool calls. We propose

Pearl

(

P

redictive

E

mbedding

A

lignment for

R

easoning in

L

atent space), a JEPA-inspired framework that learns from expert tool-use trajectories entirely in the latent space, eliminating the need for explicit tool invocation at inference time. Unlike reconstruction-based latent reasoning methods, which autoregressively generate latent tokens and suffer from training–inference mismatch and limited support for multi-step tool use,

Pearl

directly learns predictive embeddings from multimodal trajectories while preserving the standard vision-language generation pipeline: it is model-agnostic, simple to train, and naturally supports trajectories with multiple tool calls. Experiments across multiple perception benchmarks show that

Pearl

matches or outperforms standard supervised fine-tuning and reconstruction-based latent reasoning approaches. Furthermore, we provide empirical evidence that reconstruction-based methods primarily learn embeddings rather than image edits in latent space, motivating predictive embedding learning as a more principled alternative.

⟨

ℐ

0

,

𝒬

⟩

\langle\mathcal{I}_{0},\,\mathcal{Q}\rangle

ℛ

=

(

ℐ

1

,

𝒯

1

,

…

,

ℐ

N

,

𝒯

N

)

\mathcal{R}=(\mathcal{I}_{1},\mathcal{T}_{1},\dots,\mathcal{I}_{N},\mathcal{T}_{N})

⟨

ℐ

0

,

𝒬

⟩

\langle\mathcal{I}_{0},\,\mathcal{Q}\rangle

ℛ

=

(

ℐ

1

,

𝒯

1

,

…

,

ℐ

N

,

𝒯

N

)

\mathcal{R}=(\mathcal{I}_{1},\mathcal{T}_{1},\dots,\mathcal{I}_{N},\mathcal{T}_{N})

Enc

​

(

⟨

ℐ

0

,

𝒬

⟩

)

\mathrm{Enc}(\langle\mathcal{I}_{0},\mathcal{Q}\rangle)

Enc

​

(

ℛ

)

\mathrm{Enc}(\mathcal{R})

Pred

​

(

⋅

)

\mathrm{Pred}(\,\cdot\,)

[

PRED

] tokens

h

^

ℛ

\hat{h}_{\mathcal{R}}

h

ℛ

=

sg

​

[

Enc

​

(

ℛ

)

]

h_{\mathcal{R}}=\mathrm{sg}[\mathrm{Enc}(\mathcal{R})]

ℒ

VLM

\mathcal{L}_{\mathrm{VLM}}

ℒ

JEPA

=

D

​

(

h

^

ℛ

,

h

ℛ

)

\mathcal{L}_{\mathrm{JEPA}}=D(\hat{h}_{\mathcal{R}},\,h_{\mathcal{R}})

ℒ

NextLat

\mathcal{L}_{\mathrm{NextLat}}

ℒ

PEARL

=

ℒ

VLM

+

λ

​

[

ℒ

JEPA

+

ℒ

NextLat

]

\mathcal{L}_{\mathrm{PEARL}}=\mathcal{L}_{\mathrm{VLM}}+\lambda\,[\mathcal{L}_{\mathrm{JEPA}}+\mathcal{L}_{\text{NextLat}}]

Example trajectory

ℛ

\mathcal{R}

ℐ

0

\mathcal{I}_{0}

𝒬

\mathcal{Q}

:

What is the design on the purple shoes worn by the child?

input

step 1

ℐ

1

\mathcal{I}_{1}

𝒯

1

\mathcal{T}_{1}

:

Crop to the shoes of the child.

⋮

\vdots

step

N

N

ℐ

N

\mathcal{I}_{N}

𝒯

N

\mathcal{T}_{N}

:

The design is a rabbit.

Inference:

only

⟨

ℐ

0

,

𝒬

⟩

\langle\mathcal{I}_{0},\mathcal{Q}\rangle

no tools, no latent decoding

Figure 1:

Pearl

architecture.

Left:

Solid arrows
denote forward-pass dataflow; dashed arrows denote which components
contribute to each loss (no forward pass). During training, two
independent forward passes encode

⟨

ℐ

0

,

𝒬

⟩

\langle\mathcal{I}_{0},\mathcal{Q}\rangle

and the expert trajectory

ℛ

\mathcal{R}

. A tied-weights predictor maps the input encoding to
the trajectory latent space.

ℒ

JEPA

\mathcal{L}_{\mathrm{JEPA}}

aligns
the predicted embedding

h

^

ℛ

\hat{h}_{\mathcal{R}}

with the stop-gradient
target

h

ℛ

h_{\mathcal{R}}

;

ℒ

VLM

\mathcal{L}_{\mathrm{VLM}}

preserves
autoregressive text generation;

ℒ

NextLat

\mathcal{L}_{\text{NextLat}}

regularises hidden states to act as belief states.

Right:

An example trajectory of sequential visual edits

(

ℐ

1

,

…

,

ℐ

N

)

(\mathcal{I}_{1},\dots,\mathcal{I}_{N})

interleaved with reasoning
text

(

𝒯

1

,

…

,

𝒯

N

)

(\mathcal{T}_{1},\dots,\mathcal{T}_{N})

.

1

Introduction

Recent work in multimodal reasoning has explored augmenting vision-language models (VLMs) with external tools (e.g., cropping, object detection, depth estimation) to improve grounded reasoning

(Su

et al.

,

2024

; Wu

et al.

,

2025b

; Su

et al.

,

2025

)

. By enabling models to iteratively manipulate visual inputs, these approaches allow VLMs to “think with images” rather than relying solely on textual reasoning.

While interacting with expert tools to edit images is an effective strategy for grounding LLMs in visual context

(Yue

et al.

,

2024

; Hao

et al.

,

2025

)

, tool-based approaches introduce practical and conceptual challenges. First, invoking external tools incurs substantial inference-time latency and compute overhead

(Nichols

et al.

,

2025

; Wu

et al.

,

2025a

)

. Second, learning to correctly select and parameterize tools requires specialized supervision

(Wu

et al.

,

2025b

; Su

et al.

,

2024

)

, and even after such training, erroneous calls waste inference compute and pollute the context with irrelevant information. Third, most approaches assume a homogeneous tool set, whereas handling the diverse tools exposed by complex MCP servers requires advanced planning and instruction-following

(Wu

et al.

,

2025a

)

.

A promising direction replaces explicit tool use with latent
reasoning, where models operate in a continuous embedding space
instead of generating discrete intermediate outputs

(Hao

et al.

,

2024

; Tan

et al.

,

2025

; Gozeten

et al.

,

2026

)

. Prior work has explored

reconstruction-based

latent reasoning, in which models autoregressively generate latent
tokens intended to represent intermediate visual transformations

(Li

et al.

,

2025

; Yang

et al.

,

2025

; Gu

et al.

,

2025

)

. Borrowing from
Coconut

(Hao

et al.

,

2024

)

, these methods supervise continuous latent
tokens with a reconstruction objective against the outputs of visual
tools, while preserving the standard transformer architecture (see
Figure

5

for an overview). Despite their appeal, these
approaches suffer from two fundamental limitations. First, they
exhibit a training–inference mismatch: during training, models are
supervised with as many latent tokens as there are image patch tokens
in the tool output, yet at inference only a small, fixed number of
latent tokens is decoded, often without improving, and sometimes
degrading, performance

(Yang

et al.

,

2025

; Li

et al.

,

2025

)

. Second, they are typically confined to
single-step transformations, failing to support multi-step reasoning
over sequences of tool operations. These observations suggest that
reconstruction-based methods primarily learn useful

embeddings

rather than genuinely simulating visual transformations in latent
space.

In this work, we propose

Pearl

(

P

redictive

E

mbedding

A

lignment for

R

easoning in

L

atent space), a JEPA-inspired framework that learns
predictive representations from expert tool-use trajectories. Rather
than autoregressively generating latent tokens,

Pearl

predicts trajectory embeddings from an image–question pair (see
Figure

1

, right), allowing the model to internalize the
effects of tool use without explicit tool invocation. The framework
operates entirely in latent space during training, avoids
training–inference mismatch, and supports multi-step reasoning over
trajectories with multiple tool calls. We instantiate

Pearl

by jointly optimizing a standard vision–language generation objective
with a predictive embedding objective over interleaved multimodal
trajectories (see Figure

1

, left), enabling the model to
retain its text generation capabilities while learning both the

effects

and

sequencing

of task-relevant transformations.

We evaluate

Pearl

across a range of multimodal reasoning
benchmarks, including settings with single and multiple tool
calls.

Pearl

consistently matches or outperforms supervised
fine-tuning and reconstruction-based latent reasoning
approaches. Moreover, our analysis shows that reconstruction-based
methods primarily learn embeddings rather than performing genuine
latent “imagination”, supporting predictive embedding learning as a
more principled alternative.

2

Related Work

Tool-augmented Multi-modal Reasoning.

A prominent line of work augments vision-language models with external visual tools to improve grounded reasoning

(Su

et al.

,

2024

; Huang

et al.

,

2025b

; Zheng

et al.

,

2025

)

. These approaches enable models to iteratively manipulate images through operations such as cropping, object detection, or spatial transformations, effectively allowing them to “think with images” rather than relying solely on textual reasoning. More advanced systems further integrate specialized tools such as depth estimation or multi-step visual editing pipelines, often interleaving tool execution with chain-of-thought reasoning. To determine when and how to invoke tools, early methods rely on supervised fine-tuning with expert trajectories

(Su

et al.

,

2025

; Chung

et al.

,

2026

)

while more recent approaches leverage reinforcement learning to acquire tool-use policies and support multi-step reasoning

(Zheng

et al.

,

2025

; Su

et al.

,

2024

; Geng

et al.

,

2026

)

.
Despite their effectiveness, these methods introduce significant practical challenges: tool invocation incurs substantial inference-time overhead, requires specialized supervision for correct tool selection and parameterization, and remains brittle to errors that can propagate through the reasoning process. Furthermore, many approaches assume a fixed or homogeneous tool set, limiting scalability to diverse or dynamically evolving tool environments. We instead explore latent visual reasoning, where the model internalizes the effects of tool use directly in representation space, eliminating the need for explicit tool invocation at inference time.

Reconstruction-based Latent Reasoning.

Existing work on
latent reasoning in multimodal models autoregressively generates
latent tokens under a reconstruction objective, aiming to “imagine”
intermediate image edits. Concretely, these latent tokens are trained
to reconstruct tool outputs, borrowing from approaches such as
Coconut

(Hao

et al.

,

2024

)

in the text domain. This requires models to
switch between latent and discrete tokens at inference, a complication
that

Pearl

avoids. However, this line of work is largely
confined to single tool calls within a reasoning trajectory

(Li

et al.

,

2025

; Yang

et al.

,

2025

; Gu

et al.

,

2025

)

. COVT

(Qin

et al.

,

2025

)

extends this setting to multiple tool calls, but applies a fixed
sequence of operations regardless of the input query. This design
imposes two key limitations. First, COVT avoids dynamic planning by
restricting tool calls to parameter-free operations, precluding
actions such as cropping that require query-specific
arguments. Second, all tool calls are applied directly to the original
image

ℐ

0

\mathcal{I}_{0}

rather than to the outputs of preceding
operations, resulting in a shallow tree of independent branches rather
than a chain of dependent steps. In contrast,

Pearl

avoids
autoregressive latent generation by predicting a single embedding of
the full expert trajectory, naturally supporting multi-step tool use
as the prediction target encodes the entire trajectory

ℛ

\mathcal{R}

rather than a single transformation.

Joint Embedding Predictive Architectures for Language Models.

The Joint Embedding Predictive Architecture (JEPA) is a self-supervised learning framework that trains a model to predict the embedding of one view of the data from another, rather than reconstructing raw inputs. It has shown promise as a pre-training objective for multimodal models with V-JEPA2

(Assran

et al.

,

2025

)

and VL-JEPA

(Chen

et al.

,

2026

)

achieving competitive performance; however, both require substantial pre-training compute to match the performance of current state-of-the-art models. In the textual domain, LLM-JEPA

(Huang

et al.

,

2025a

)

adapts
JEPA to fine-tune off-the-shelf language models, avoiding the need to train from scratch. Concretely, LLM-JEPA minimises a distance between the encoded representations of paired text and code, encouraging the model to develop modality-agnostic representations conducive to text-to-code generation. Similarly, V-JEPA2 aligns embeddings from video frames and
robotic actions with predicted future frames for video model pre-training.

Pearl

adapts the JEPA objective to fine-tune an off-the-shelf VLM from expert multimodal tool-use trajectories, treating the image-question pair and the full reasoning trajectory as two views of the same problem. This allows

Pearl

to internalize the effect of sequential visual tool use in the latent space, without incurring the prohibitive compute of pre-training or departing from the standard image-text-to-text inference pipeline.

3

PEARL: Predictive Latent Reasoning

3.1

Problem Formulation and Overview

We consider a training setting where each example consists of an
image-question pair

𝐱

=

⟨

ℐ

0

,

𝒬

⟩

\mathbf{x}=\langle\mathcal{I}_{0},\mathcal{Q}\rangle

and an expert multimodal reasoning trajectory

ℛ

=

(

ℐ

1

,

𝒯

1

,

ℐ

2

,

𝒯

2

,

…

,

ℐ

N

,

𝒯

N

)

,

\mathcal{R}=(\mathcal{I}_{1},\mathcal{T}_{1},\mathcal{I}_{2},\mathcal{T}_{2},\dots,\mathcal{I}_{N},\mathcal{T}_{N}),

where each

ℐ

i

\mathcal{I}_{i}

is an intermediate image produced by an
expert visual tool (e.g., crop, highlight, spatial transformation) and
each

𝒯

i

\mathcal{T}_{i}

is the associated reasoning text, with the final
step containing the answer to

𝒬

\mathcal{Q}

(see
Figure

1

).

Our goal is to train a VLM that benefits from such tool-use trajectories

without

invoking tools at inference time. In contrast to prior
reconstruction-based approaches, which autoregressively generate latent
tokens intended to reconstruct intermediate visual edits

(Li

et al.

,

2025

; Yang

et al.

,

2025

; Gu

et al.

,

2025

)

,

Pearl

(

P

redictive

E

mbedding

A

lignment for

R

easoning in

L

atent space) directly predicts a latent
representation of the full trajectory from the original image-question
pair, preserving the standard VLM inference pipeline while internalizing
information from tool-based reasoning during training.

Concretely,

Pearl

encodes

𝐱

\mathbf{x}

and

ℛ

\mathcal{R}

independently and trains a lightweight predictor to anticipate the
trajectory embedding from the input alone. Intuitively, the predictor asks: given only the image and question,
can the model anticipate what the expert tool-use trajectory would look like
in latent space? This design has three
advantages. First, it avoids explicit tool invocation at inference.
Second, it avoids the training–inference mismatch of reconstruction-based
methods, where models are trained with many latent tokens but decode only
a small fixed number at test time. Third, because the prediction target
encodes the

entire

multimodal trajectory rather than a single
image edit,

Pearl

naturally supports multiple tool calls.

3.2

Trajectory Encoding

We instantiate both

Enc

​

(

𝐱

)

\mathrm{Enc}(\mathbf{x})

and

Enc

​

(

ℛ

)

\mathrm{Enc}(\mathcal{R})

using the hidden states of an off-the-shelf autoregressive VLM, serializing the two
views as follows (see also Figure

1

):

•

the

input view

consists of the original image-question pair

⟨

ℐ

0

,

𝒬

⟩

\langle\mathcal{I}_{0},\mathcal{Q}\rangle

;

•

the

trajectory view

consists of the interleaved sequence

(

ℐ

1

,

𝒯

1

,

…

,

ℐ

N

,

𝒯

N

)

(\mathcal{I}_{1},\mathcal{T}_{1},\ldots,\mathcal{I}_{N},\mathcal{T}_{N})

.

For each view, we run a forward pass through the VLM and take the final hidden
state of the last token as its sequence representation, following prior work on
JEPA-style fine-tuning of decoder-only language
models

(Huang

et al.

,

2025a

)

:

h

𝐱

=

Enc

​

(

𝐱

)

,

h

ℛ

=

Enc

​

(

ℛ

)

.

h_{\mathbf{x}}=\mathrm{Enc}(\mathbf{x}),\qquad h_{\mathcal{R}}=\mathrm{Enc}(\mathcal{R}).

Using separate forward passes avoids cross-view information leakage and keeps
the method architecture-agnostic, at the cost of additional training-time
compute. This overhead applies only during training; inference remains identical
to standard VLM decoding. A lightweight predictor network then takes
the input encoding

h

𝐱

h_{\mathbf{x}}

and produces a predicted version of
the trajectory embedding.

3.3

Latent Trajectory Predictor

To map the input representation

h

𝐱

h_{\mathbf{x}}

to the trajectory
latent space, we use a predictor built from the VLM itself. Following
prior tied-weights JEPA formulations

(Huang

et al.

,

2025a

)

, we append

K

K

learnable
special tokens

[PRED]

to the serialized input

𝐱

\mathbf{x}

,
and define the predicted trajectory representation as the hidden state
of the final predictor token:

h

^

ℛ

=

Pred

​

(

h

𝐱

)

.

\hat{h}_{\mathcal{R}}=\mathrm{Pred}(h_{\mathbf{x}}).

Intuitively, the predictor allows the model to perform additional
nonlinear computation over the image-question representation before
producing the target latent (see visualization in Figure

1

). When

K

=

0

K=0

, the predictor reduces to
the identity map. In practice, using predictor tokens lets us reuse
the VLM’s existing self-attention stack rather than introducing a
separate MLP or auxiliary transformer, thereby keeping the method
simple and parameter-efficient.

3.4

Predictive Embedding Objective

Our central training signal is a JEPA-style predictive embedding loss that aligns
the predicted latent

h

^

ℛ

\hat{h}_{\mathcal{R}}

with the encoded expert trajectory

h

ℛ

h_{\mathcal{R}}

. We define

ℒ

JEPA

=

D

​

(

h

^

ℛ

,

sg

​

[

h

ℛ

]

)

,

\mathcal{L}_{\rm JEPA}=D\!\left(\hat{h}_{\mathcal{R}},\mathrm{sg}[h_{\mathcal{R}}]\right),

(1)

where

D

​

(

⋅

,

⋅

)

D(\cdot,\cdot)

is a distance function and

sg

​

[

⋅

]

\mathrm{sg}[\cdot]

denotes
stop-gradient. In our experiments, we use SmoothL1 loss for

D

D

.

This objective encourages the model to learn a compact representation of the

effect

of expert tool use and multimodal reasoning, rather than explicitly
reconstructing intermediate image edits. In this sense, PEARL learns predictive
trajectory embeddings rather than latent image generation (see Figure

1

).

3.5

Next-Latent Prediction

A potential limitation of using the final hidden state of a decoder as
a sequence representation is that it may not reliably summarize all
relevant preceding context. To encourage hidden states to behave as
predictive summary states, we add a next-latent prediction objective
inspired by recent work on latent dynamics in transformers

(Teoh

et al.

,

2025

)

.

Let

h

t

h_{t}

denote the hidden state at time step

t

t

within the
serialized trajectory. A lightweight latent predictor is trained to
forecast future hidden states

h

^

t

+

i

\hat{h}_{t+i}

from the current state

h

t

h_{t}

, for a prediction horizon

d

d

. We optimize

ℒ

NextLat

=

𝔼

t

​

[

1

d

​

∑

i

=

1

d

SmoothL1Loss

​

(

sg

​

[

h

t

+

i

]

,

h

^

t

+

i

)

]

.

\mathcal{L}_{\text{NextLat}}=\mathbb{E}_{t}\left[\frac{1}{d}\sum_{i=1}^{d}\mathrm{SmoothL1Loss}\left(\mathrm{sg}[h_{t+i}],\hat{h}_{t+i}\right)\right].

(2)

This objective encourages hidden states to be informative about future trajectory evolution, making them better suited for sequence-level latent alignment. Specifically,

Teoh

et al.

(

2025

)

show that optimizing the hidden state transitions as in Equation (

2

) causes them to converge to

belief states

, which

Kaelbling

et al.

(

1998

)

define as sufficient statistics of the past history. We view this term as a regularizer that improves the quality of the learned latent representations, rather than as a separate reasoning mechanism.

3.6

Autoregressive Generation Objective

In addition to latent alignment, we retain the standard VLM training
objective over the textual portions of the expert
trajectory. Given the interleaved multimodal context, the
model is trained to autoregressively predict each token in the textual
segments

𝒯

1

,

…

,

𝒯

N

\mathcal{T}_{1},\dots,\mathcal{T}_{N}

:

ℒ

VLM

=

−

∑

n

=

1

N

∑

t

=

1

|

𝒯

n

|

log

⁡

p

θ

​

(

𝒯

n

(

t

)

∣

ℐ

0

,

𝒬

,

ℐ

1

,

𝒯

1

,

…

,

ℐ

n

,

𝒯

n

(

<

t

)

)

.

\mathcal{L}_{\rm VLM}=-\sum_{n=1}^{N}\sum_{t=1}^{|\mathcal{T}_{n}|}\log p_{\theta}\!\left(\mathcal{T}_{n}^{(t)}\mid\mathcal{I}_{0},\mathcal{Q},\mathcal{I}_{1},\mathcal{T}_{1},\dots,\mathcal{I}_{n},\mathcal{T}_{n}^{(<t)}\right).

(3)

Here

𝒯

n

(

t

)

\mathcal{T}_{n}^{(t)}

denotes the

t

t

-th token of the

n

n

-th
reasoning step,

𝒯

n

(

<

t

)

\mathcal{T}_{n}^{(<t)}

denotes all preceding tokens
within that step, and the conditioning context includes all prior
image-text pairs

(

ℐ

1

,

𝒯

1

,

…

,

ℐ

n

−

1

,

𝒯

n

−

1

)

(\mathcal{I}_{1},\mathcal{T}_{1},\dots,\mathcal{I}_{n-1},\mathcal{T}_{n-1})

as well as the original input

⟨

ℐ

0

,

𝒬

⟩

\langle\mathcal{I}_{0},\mathcal{Q}\rangle

. This term ensures that
PEARL preserves the VLM’s standard text generation capability, which
is necessary for producing final answers at test time.

3.7

Training Objective

We jointly optimize the autoregressive generation objective, the predictive
embedding objective, and the next-latent regularizer:

ℒ

PEARL

=

ℒ

VLM

+

λ

​

[

ℒ

JEPA

+

ℒ

NextLat

]

,

\mathcal{L}_{\rm PEARL}=\mathcal{L}_{\rm VLM}+\lambda\,[\mathcal{L}_{\rm JEPA}+\mathcal{L}_{\text{NextLat}}],

(4)

where

λ

\lambda

jointly controls the contribution of both latent
objectives relative to the generation loss, reflecting the view that

ℒ

JEPA

\mathcal{L}_{\rm JEPA}

and

ℒ

NextLat

\mathcal{L}_{\text{NextLat}}

together
constitute a single latent learning signal (see Figure

1

).

These three terms play complementary roles.

ℒ

VLM

\mathcal{L}_{\rm VLM}

preserves the model’s ability to generate answers in the discrete
token space.

ℒ

JEPA

\mathcal{L}_{\rm JEPA}

teaches the model to predict
a latent representation of expert multimodal reasoning from the
original image-question pair.

ℒ

NextLat

\mathcal{L}_{\text{NextLat}}

encourages hidden states to act as
belief states (i.e., sufficient summaries of past context), making them
more informative encoding targets for

ℒ

JEPA

\mathcal{L}_{\rm JEPA}

.

3.8

Inference

At inference time, PEARL requires only the original image-question
pair

⟨

ℐ

0

,

𝒬

⟩

\langle\mathcal{I}_{0},\mathcal{Q}\rangle

, and answers using
the standard generation pipeline of the underlying VLM. It does

not

invoke external tools, does

not

generate
intermediate edited images, does

not

use

[PRED]

tokens, and does

not

autoregressively decode latent reasoning
tokens. The cost of learning from tool use is shifted entirely to
training time, preserving simple and efficient inference.

4

Experimental Setting

Training Regimes.

To demonstrate the effectiveness of

Pearl

at learning from
expert tool-use trajectories, we finetune models across three settings:
(i) single-type, single tool call per trajectory;
(ii) multiple-type, single tool call per trajectory; and
(iii) single-type, multiple tool calls per trajectory.
We leave the multiple-type, multiple tool call setting to future work,
as no open-source training data currently exists for this combination.

For setting (i), we use the data from LVR

(Li

et al.

,

2025

)

,
which provides regions of interest used to crop the input image

ℐ

0

\mathcal{I}_{0}

, forming the trajectory

ℛ

\mathcal{R}

.
For setting (ii), we use the ThinkMorph
dataset

(Gu

et al.

,

2025

)

, which
contains four equal-sized subsets corresponding to different tool types:
bounding boxes over regions of interest, highlights over charts, jigsaw
puzzle reconstructions, and spatial navigation paths over maze images.
For setting (iii), we use the PixelReasoner
dataset

(Su

et al.

,

2024

)

, where each trajectory

ℛ

\mathcal{R}

contains up to three sequential crops of

ℐ

0

\mathcal{I}_{0}

. Dataset
statistics and examples are provided in
Appendix

B

.

Evaluation Benchmarks.

Following previous work (e.g.,

Li

et al.

2025

)
we evaluate

Pearl

on a suite of perception intensive visual
question answering (VQA;

Antol

et al.

2015

) benchmarks. These
include V*

(Wu and Xie,

2023

)

, which tests models’ ability to perform
visual search for objects and their attributes (V*

DA

) and to
identify relative positions of objects (V*

RP

). We further
evaluate on five subsets of the Blink benchmark

(Fu

et al.

,

2025

)

:
Counting, IQ, Jigsaw, Spatial Relation, and Relative
Reflectance. Finally, we include MMVP

(Tong

et al.

,

2024

)

, which probes
perceptual robustness using image pairs that CLIP treats as similar
despite clear visual differences. All benchmarks are formulated as
multiple-choice tasks, enabling straightforward answer parsing (see
Appendix

B

for details).

Comparison Models.

Our primary comparisons use Qwen2.5-VL-7B-Instruct

(Team,

2025

)

, enabling direct head-to-head
evaluation against all reconstruction-based baselines. To further demonstrate

Pearl

’s model-agnostic nature, we also report
results for the smaller Qwen2.5-VL-3B-Instruct variant and the 4B variant of Qwen3-VL

(Bai

et al.

,

2025

)

.

For the single-type, single tool call setting, we compare against
LVR

(Li

et al.

,

2025

)

, using their released
HuggingFace checkpoint, with 4 latent tokens (aka 4 steps) which the
authors note yields the best overall quality. We also compare
against CoVT

(Qin

et al.

,

2025

)

, reporting results directly from the
original paper. LVR achieves the strongest performance among
reconstruction-based latent reasoning methods (see
Figure

5

for an illustration).

For the multiple-type, single tool call setting, we compare

Pearl

against a LoRA-finetuned variant trained on the
ThinkMorph data

(Gu

et al.

,

2025

)

.
We do not compare against the original ThinkMorph model, as it relies
on explicit intermediate image generation at inference, making it
incomparable with latent reasoning methods.

1

1

1

We were also
unable to reproduce the original ThinkMorph results, as the released
checkpoint and evaluation scripts were not available in a complete
form at the time of submission.

For the single-type, multiple tool
call setting, we compare directly against
PixelReasoner’s (

2024

) released model, which
invokes tools explicitly at inference.

Pearl

is finetuned with LoRA

(Hu

et al.

,

2021

)

adapters (rank

r

=

64

r=64

and

α

=

128

\alpha=128

).
Across settings, we include a
LoRA SFT baseline trained on the same data as

Pearl

and the
instruction-tuned model without fine-tuning as a zero-shot
baseline. Hyperparameter settings for

Pearl

are provided in
Appendix

A

.

5

Results

5.1

How Does

Pearl

Compare to Reconstruction-based Methods?

Model

V

∗

V

D

​

A

∗

{}^{*}_{DA}

V

R

​

P

∗

{}^{*}_{RP}

MMVP

Counting

IQ

Jigsaw

Rel. Ref

Spatial Rel

No fine-tuning

Qwen2.5-VL-7B-Instruct

78.5

81.7

73.7

66.7

66.7

26.0

52.0

38.8

87.4

Single-type, single tool call

CoVT

(Qin

et al.

,

2025

)

78.0

—

—

58.7

—

—

—

—

—

LVR

(Li

et al.

,

2025

)

(4 steps)

80.1

85.2

73.7

72.0

68.3

26.0

51.3

41.0

89.5

SFT (LVR data)

79.1

82.6

73.7

65.7

67.5

26.7

45.3

33.6

88.8

Pearl

(LVR data)

81.5

86.1

74.5

73.5

68.3

28.2

53.1

39.6

89.5

Multiple-type, single tool call

SFT (ThinkMorph data)

42.4

58.3

18.4

36.7

38.3

16.7

22.0

38.8

60.1

Pearl

(ThinkMorph data)

73.8

76.5

69.7

75.3

65.0

26.0

53.3

46.3

88.8

Single-type, multiple tool calls

PixelReasoner

(Su

et al.

,

2024

)

80.1

81.7

77.6

67.0

66.7

25.3

52.7

42.5

88.1

Pearl

(PixelReasoner data)

79.1

81.7

75.0

70.0

70.0

28.7

53.3

40.3

89.5

Table 1:

Results for Qwen2.5-VL-7B-Instruct across all training settings.

Pearl

requires no
tool calls at inference time.

Bold

denotes best result
per block.

Table

1

compares

Pearl

against various
Qwen2.5-VL-7B-Instruct baselines and comparison systems (across three training settings).
As can be seen,

Pearl

consistently matches or outperforms
its respective baselines while requiring no tool calls at inference,
an advantage none of the reconstruction-based or tool-augmented methods
share.

Single-type, single tool call.

Pearl

trained on LVR data outperforms both the SFT baseline
and LVR

(Li

et al.

,

2025

)

(4 steps) on V

∗

(81.5
vs. 79.1 and 80.1) and MMVP (73.5 vs. 65.7 and 72.0), while matching
LVR on Spatial Rel (89.5) and improving on Jigsaw (53.1 vs. 51.3).
Notably, LVR finetunes the entire decoder whereas

Pearl

uses
only LoRA adapters, making these gains more parameter-efficient.
CoVT

(Qin

et al.

,

2025

)

underperforms even the zero-shot baseline on
MMVP (58.7 vs. 66.7), suggesting its fixed-sequence design is poorly
suited to this benchmark.

Multiple-type, single tool call.

The ThinkMorph results are the most striking in the table.

Pearl

outperforms the SFT baseline by over 31 points on V

∗

(73.8 vs. 42.4) and more than doubles it on MMVP (75.3 vs. 36.7).
The SFT baseline collapses under the heterogeneity of four qualitatively
different tool types, whereas

Pearl

’s trajectory-level embedding
target is agnostic to tool type, explaining its robustness across the
full ThinkMorph benchmark.

Single-type, multiple tool calls.

Pearl

is competitive with PixelReasoner

(Su

et al.

,

2024

)

,
which explicitly invokes tools at inference time.

Pearl

outperforms it on MMVP (70.0 vs. 67.0), Counting (70.0 vs. 66.7),
Jigsaw (53.3 vs. 52.7), IQ (28.7 vs. 25.3), and Spatial Rel (89.5
vs. 88.1), while PixelReasoner leads on V

R

​

P

∗

{}^{*}_{RP}

(77.6 vs. 75.0)
and Rel. Ref (42.5 vs. 40.3). The fact that

Pearl

matches an
inference-time tool-use system while operating as a standard
image-to-text model demonstrates that the tool-use signal can be
effectively internalized during training through predictive embedding
alignment.

Ablations in Appendix

A

confirm that encouraging hidden states to act as belief states
meaningfully improves the quality of the learned trajectory embeddings. Figure

4

provides further support:
t-SNE visualizations show that

Pearl

induces coherent clusters that
align the two views (

⟨

ℐ

0

,

𝒬

⟩

\langle\mathcal{I}_{0},\mathcal{Q}\rangle

and

ℛ

\mathcal{R}

) across tasks, whereas SFT produces fragmented clusters,
confirming that predictive embedding alignment learns more structured
representations.

5.2

What is the Effect of Training Regime on

Pearl

?

Although
no single training regime dominates uniformly across all benchmarks,
clear patterns emerge. The LVR regime (single-type, single tool
call) is strongest on

visual search

tasks, yielding the highest scores
on V

∗

(81.5) and V

D

​

A

∗

{}^{*}_{DA}

(86.1). The PixelReasoner regime
(single-type, multiple tool calls) performs best on tasks requiring

counting and spatial reasoning

, leading on Counting (70.0) and matching
the best result on Spatial Rel (89.5). The ThinkMorph regime
(multiple-type, single tool call) stands out on

perceptual robustness

benchmarks, leading clearly on MMVP (75.3) and Rel. Ref (46.3),
suggesting that exposure to diverse tool types improves fine-grained
perceptual discrimination.
Taken together, these results indicate that the three regimes are
complementary rather than competing.A natural direction for future work is a combined training strategy
that draws on all three regimes simultaneously, which we would expect
to yield stronger across-the-board performance.

Model

V

∗

V

D

​

A

∗

{}^{*}_{DA}

V

R

​

P

∗

{}^{*}_{RP}

MMVP

Counting

IQ

Jigsaw

Rel. Ref

Spatial Rel

Qwen2.5-VL-3B-Instruct

No fine-tuning

56.0

53.0

60.5

59.3

65.8

26.0

45.3

44.8

80.4

LVR

(Li

et al.

,

2025

)

(4 steps)

64.9

69.6

60.5

54.7

—

29.3

52.7

—

—

Pearl

(LVR data)

73.8

82.6

60.5

68.7

66.7

29.3

51.3

41.8

84.6

Pearl

(PixelReasoner data)

69.6

78.3

56.6

63.7

67.5

30.7

49.3

43.3

81.8

Pearl

(ThinkMorph data)

62.8

69.6

52.6

60.0

61.7

29.3

43.3

35.1

78.3

Qwen3-VL-4B-Instruct

No fine-tuning

81.2

86.1

73.7

75.7

65.8

24.0

68.0

62.7

83.9

Pearl

(LVR data)

81.7

85.2

76.3

80.0

67.5

27.3

70.0

53.7

85.3

Pearl

(PixelReasoner data)

79.6

82.6

75.0

77.3

70.8

25.3

68.0

57.5

87.4

Pearl

(ThinkMorph data)

75.4

81.7

65.8

76.3

66.7

26.7

75.3

66.4

81.8

Table 2:

Results for smaller model variants (Qwen2.5-VL-3B-Instruct and Qwen3-VL-4B-Instruct) across three training regimes.

Bold

denotes best result per block.

5.3

Does

Pearl

Generalise Across Model Sizes and Architectures?

Table

2

reports results for smaller model variants,
demonstrating that

Pearl

’s gains are not specific to the 7B
scale or to the Qwen2.5 architecture.

Pearl

trained on LVR data substantially outperforms the 3B LVR
baseline across nearly all benchmarks — most strikingly on V

∗

(73.8
vs. 64.9) and MMVP (68.7 vs. 54.7) — despite using only LoRA
adapters. This mirrors the pattern observed at 7B and confirms that
predictive embedding learning scales down gracefully. The PixelReasoner-trained
variant performs slightly lower overall but remains competitive on
Counting and Rel. Ref, consistent with the regime-specific patterns
observed in Table

1

.

The zero-shot Qwen3-VL-4B baseline is already strong, particularly on
Jigsaw (68.0) and Rel. Ref (62.7), which substantially exceed the
corresponding 7B zero-shot scores, reflecting the stronger
perceptual capabilities of the Qwen3 architecture.

Pearl

trained
on LVR data improves further on V

∗

(81.7 vs. 81.2) and MMVP (80.0
vs. 75.7), while the PixelReasoner-trained variant gains on
Spatial Rel (87.4 vs. 83.9). Both variants show some regression on
Rel. Ref relative to the zero-shot baseline, which we leave to future
investigation.

For the 3B variant,

Pearl

trained on ThinkMorph underperforms
the LVR-trained variant across most benchmarks (V

∗

: 62.8 vs. 73.8;
MMVP: 60.0 vs. 68.7). This is likely due to two compounding factors:
the diverse tool-type signal may require greater model capacity, and
ThinkMorph’s verbose, open-ended reasoning steps introduce a
training–test format mismatch that smaller models struggle to overcome
when producing multiple-choice answers. By contrast, the 4B Qwen3-VL
variant benefits strongly from the ThinkMorph regime, achieving 75.3
on Jigsaw (vs. 68.0 zero-shot) and 66.4 on Rel. Ref
(vs. 62.7 zero-shot), surpassing both the LVR- and PixelReasoner-trained
variants on these benchmarks — consistent with the 7B finding that
diverse tool exposure improves perceptual discrimination.

Across both a smaller and a newer model architecture,

Pearl

consistently matches or improves over its respective fine-tuning
baselines without any architecture-specific modifications, confirming
its model-agnostic nature. The
complementary strengths of the three training regimes identified at 7B (visual search (LVR), spatial and counting abilities (PixelReasoner), and
perceptual robustness (ThinkMorph)) generalise across scales, although the
ThinkMorph gains appear sensitive to base model capacity.

5.4

Do Reconstruction-based Methods Actually “Imagine” Images?

A central motivation for

Pearl

is the observation that
reconstruction-based latent reasoning methods

(Li

et al.

,

2025

; Yang

et al.

,

2025

; Qin

et al.

,

2025

)

may not be doing
what they claim. These methods assume that autoregressively
generating latent tokens allows a model to “imagine” intermediate
image edits in the latent space, and that more tokens should
correspond to a more complete imagined transformation.

Figure 2:

Cumulative distribution function (CDF) of the number of
latent tokens per training example (x-axis) over a sample of

∼

\sim

19k examples used to train LVR.

Figure

2

shows that over 75% of the edited
images used to supervise LVR contain more than 8 latent tokens during
training, as a direct consequence of the token count scaling with the
number of image patch tokens in each example, yet LVR fixes this to
just 4 or 8 tokens at inference. Figure

3

further reveals that model quality does not improve as the number of
latent tokens increases, and in some cases slightly degrades, with a
weakly negative correlation across BLINK and MMVP. In fact, using just
1 or 2 latent tokens achieves parity with much higher token counts.
This training–inference mismatch, combined with the insensitivity of
performance to token count, suggests that reconstruction-based methods
are not genuinely simulating visual transformations in latent
space. Instead, they appear to learn useful

embeddings

: compact
representations that improve answer quality regardless of how many
latent tokens are decoded. This finding directly motivates

Pearl

: if reconstruction-based methods are learning
embeddings anyway, it is more principled to learn these
directly via a predictive objective, without the added complexity of
autoregressive latent generation and the practical burden of switching
between continuous and discrete tokens at inference.

1

2

4

8

16

32

64

128

53.5

53.5

54

54

54.5

54.5

55

55

55.5

55.5

56

56

56.5

56.5

57

57

57.5

57.5

r

=

−

0.35

r\!=\!-0.35

,

R

2

=

0.12

R^{2}\!=\!0.12

steps (

log

2

\log_{2}

)

accuracy (%)

BLINK

1

2

4

8

16

32

64

128

77

77

78

78

79

79

80

80

81

81

82

82

83

83

r

=

0.72

r\!=\!0.72

,

R

2

=

0.51

R^{2}\!=\!0.51

steps (

log

2

\log_{2}

)

V

∗

1

2

4

8

16

32

64

128

70

70

70.5

70.5

71

71

71.5

71.5

72

72

72.5

72.5

73

73

73.5

73.5

r

=

−

0.56

r\!=\!-0.56

,

R

2

=

0.32

R^{2}\!=\!0.32

steps (

log

2

\log_{2}

)

MMVP

1

2

4

8

16

32

64

128

62

62

62.5

62.5

63

63

63.5

63.5

64

64

64.5

64.5

65

65

r

=

−

0.12

r\!=\!-0.12

,

R

2

=

0.02

R^{2}\!=\!0.02

steps (

log

2

\log_{2}

)

Average

Figure 3:

Correlation between the number of reasoning steps and accuracy across
BLINK (

n

=

697

n{=}697

), V

∗

(

n

=

191

n{=}191

), MMVP (

n

=

300

n{=}300

), and on
average (

n

=

1

,

188

n{=}1{,}188

). The x-axis uses a

log

2

\log_{2}

scale; dashed lines show
the log-linear trend. Near-zero

r

r

and

R

2

R^{2}

values confirm that embedding
quality is stable across reasoning steps. The red line is

Pearl

trained on LVR.

6

Conclusion

We presented

Pearl

, a JEPA-inspired framework that learns
from expert tool-use trajectories in the latent space without
requiring explicit tool invocation at inference. Rather than
reconstructing intermediate image edits autoregressively,

Pearl

directly predicts a trajectory-level embedding from the
image-question pair, preserving the standard VLM inference
pipeline. Across three training regimes and multiple perception
benchmarks,

Pearl

consistently matches or outperforms
reconstruction-based methods and SFT baselines using only LoRA
adapters, with gains that generalise across model sizes and
architectures. Our analysis further challenges the premise of
reconstruction-based latent reasoning: performance is largely
insensitive to the number of latent tokens decoded at inference,
suggesting these methods learn useful embeddings rather than genuinely
imagining image edits. A natural direction for future work is a
combined training strategy that draws on all three regimes
simultaneously, as well as extending

Pearl

to settings with
diverse, multi-step tool use and explicit latent planning at inference
(see Appendix

C

for further discussion).

Ethics Statement

This work presents

Pearl

, a framework for training
vision-language models to internalize the effects of visual tool use
in the latent space. We discuss the ethical considerations most
relevant to this research.

Intended Use and Misuse.

Pearl

is designed to improve the efficiency and accuracy of
multimodal reasoning in VLMs for perception-intensive tasks. As with
any method that improves the capability of language models, there is
potential for misuse in applications that generate misleading visual
interpretations or automate harmful decision-making. We encourage
practitioners to apply appropriate safeguards when deploying systems
built on this work in high-stakes settings.

Data and Bias.

Our experiments rely on publicly available datasets (LVR, ThinkMorph,
PixelReasoner) and pre-trained models (Qwen2.5-VL, Qwen3-VL). Any
biases present in these data sources or base models may be inherited
or amplified by

Pearl

. We did not conduct a systematic bias
audit and caution against deployment in sensitive domains without
further evaluation.

Environmental Cost.

Training

Pearl

requires two forward passes per example,
roughly doubling compute relative to standard fine-tuning. All
experiments were conducted on H100 and H200 GPUs. We partially
mitigate this cost by using LoRA adapters rather than full fine-tuning,
and by training for a limited number of steps or epochs per regime.

Broader Impact.

By eliminating the need for explicit tool invocation at inference,

Pearl

reduces the latency and resource cost of deploying
tool-augmented VLMs, which may make capable multimodal reasoning
more accessible. We release model weights and code to support
reproducibility and further research.

References

S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh (2015)

Vqa: visual question answering

.

In

Proceedings of the IEEE international conference on computer vision

,

pp. 2425–2433

.

Cited by:

§4

.

M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, Mojtaba, Komeili, M. Muckley, A. Rizvi, C. Roberts, K. Sinha, A. Zholus, S. Arnaud, A. Gejji, A. Martin, F. R. Hogan, D. Dugas, P. Bojanowski, V. Khalidov, P. Labatut, F. Massa, M. Szafraniec, K. Krishnakumar, Y. Li, X. Ma, S. Chandar, F. Meier, Y. LeCun, M. Rabbat, and N. Ballas (2025)

V-jepa 2: self-supervised video models enable understanding, prediction and planning

.

External Links:

2506.09985

,

Link

Cited by:

§2

.

S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, W. Ge, Z. Guo, Q. Huang, J. Huang, F. Huang, B. Hui, S. Jiang, Z. Li, M. Li, M. Li, K. Li, Z. Lin, J. Lin, X. Liu, J. Liu, C. Liu, Y. Liu, D. Liu, S. Liu, D. Lu, R. Luo, C. Lv, R. Men, L. Meng, X. Ren, X. Ren, S. Song, Y. Sun, J. Tang, J. Tu, J. Wan, P. Wang, P. Wang, Q. Wang, Y. Wang, T. Xie, Y. Xu, H. Xu, J. Xu, Z. Yang, M. Yang, J. Yang, A. Yang, B. Yu, F. Zhang, H. Zhang, X. Zhang, B. Zheng, H. Zhong, J. Zhou, F. Zhou, J. Zhou, Y. Zhu, and K. Zhu (2025)

Qwen3-vl technical report

.

External Links:

2511.21631

,

Link

Cited by:

§4

.

D. Chen, M. Shukor, T. Moutakanni, W. Chung, J. Yu, T. Kasarla, Y. Bang, A. Bolourchi, Y. LeCun, and P. Fung (2026)

VL-jepa: joint embedding predictive architecture for vision-language

.

External Links:

2512.10942

,

Link

Cited by:

§2

.

J. Chung, J. Kim, S. Kim, J. Lee, M. S. Kim, and Y. Yu (2026)

V1: learning to point visual tokens for multimodal grounded reasoning

.

External Links:

2505.18842

,

Link

Cited by:

§2

.

X. Fu, Y. Hu, B. Li, Y. Feng, H. Wang, X. Lin, D. Roth, N. A. Smith, W. Ma, and R. Krishna (2025)

BLINK: multimodal large language models can see but not perceive

.

In

Computer Vision – ECCV 2024

,

A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol (Eds.)

,

Cham

,

pp. 148–166

.

External Links:

ISBN 978-3-031-73337-6

Cited by:

§B.3

,

§4

.

X. Geng, P. Xia, Z. Zhang, X. Wang, Q. Wang, R. Ding, C. Wang, J. Wu, K. Li, Y. Zhao, H. Yin, Y. Jiang, P. Xie, F. Huang, H. Yao, Y. R. Fung, and J. Zhou (2026)

WebWatcher: breaking new frontiers of vision-language deep research agent

.

In

The Fourteenth International Conference on Learning Representations

,

External Links:

Link

Cited by:

§2

.

H. A. Gozeten, M. E. Ildiz, X. Zhang, H. Harutyunyan, A. S. Rawat, and S. Oymak (2026)

Continuous chain of thought enables parallel exploration and reasoning

.

In

The Fourteenth International Conference on Learning Representations

,

External Links:

Link

Cited by:

§1

.

J. Gu, Y. Hao, H. W. Wang, L. Li, M. Q. Shieh, Y. Choi, R. Krishna, and Y. Cheng (2025)

ThinkMorph: emergent properties in multimodal interleaved chain-of-thought reasoning

.

External Links:

2510.27492

,

Link

Cited by:

§B.2

,

§B.2

,

§1

,

§2

,

§3.1

,

§4

,

§4

.

S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y. Tian (2024)

Training large language models to reason in a continuous latent space

.

External Links:

2412.06769

,

Link

Cited by:

§1

,

§2

.

Y. Hao, J. Gu, H. W. Wang, L. Li, Z. Yang, L. Wang, and Y. Cheng (2025)

Can mllms reason in multimodality? emma: an enhanced multimodal reasoning benchmark

.

arXiv preprint arXiv:2501.05444

.

Cited by:

§1

.

E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, and W. Chen (2021)

LoRA: low-rank adaptation of large language models

.

CoRR

abs/2106.09685

.

External Links:

Link

,

2106.09685

Cited by:

§4

.

H. Huang, Y. LeCun, and R. Balestriero (2025a)

LLM-jepa: large language models meet joint embedding predictive architectures

.

External Links:

2509.14252

,

Link

Cited by:

§2

,

§3.2

,

§3.3

.

Z. Huang, Y. Ji, A. S. Rajan, Z. Cai, W. Xiao, J. Hu, and Y. J. Lee (2025b)

VisualToolAgent (VisTA): a reinforcement learning framework for visual tool selection

.

External Links:

2505.20289

,

Link

Cited by:

§2

.

L. P. Kaelbling, M. L. Littman, and A. R. Cassandra (1998)

Planning and acting in partially observable stochastic domains

.

Artificial Intelligence

101

(

1

),

pp. 99–134

.

External Links:

ISSN 0004-3702

,

Document

,

Link

Cited by:

§3.5

.

B. Li, X. Sun, J. Liu, Z. Wang, J. Wu, X. Yu, H. Chen, E. Barsoum, M. Chen, and Z. Liu (2025)

Latent visual reasoning

.

External Links:

2509.24251

,

Link

Cited by:

§A.1

,

§B.2

,

§B.2

,

§1

,

§2

,

§3.1

,

§4

,

§4

,

§4

,

§5.1

,

§5.4

,

Table 1

,

Table 2

.

D. Nichols, P. Singhania, C. Jekel, A. Bhatele, and H. Menon (2025)

Optimizing agentic language model inference via speculative tool calls

.

External Links:

2512.15834

,

Link

Cited by:

§1

.

Y. Qin, B. Wei, J. Ge, K. Kallidromitis, S. Fu, T. Darrell, and X. Wang (2025)

Chain-of-visual-thought: teaching vlms to see and think better with continuous visual tokens

.

arXiv preprint arXiv:2511.19418

.

Cited by:

§2

,

§4

,

§5.1

,

§5.4

,

Table 1

.

H. Shao, S. Qian, H. Xiao, G. Song, Z. Zong, L. Wang, Y. Liu, and H. Li (2024)

Visual cot: unleashing chain-of-thought reasoning in multi-modal language models

.

External Links:

2403.16999

Cited by:

§B.2

.

A. Su, H. Wang, W. Ren, F. Lin, and W. Chen (2024)

Pixel reasoner: incentivizing pixel-space reasoning with curiosity-driven reinforcement learning

.

arXiv preprint arXiv:2505.15966

.

Cited by:

§B.2

,

§B.2

,

§1

,

§1

,

§2

,

§4

,

§4

,

§5.1

,

Table 1

.

Z. Y. Su, L. Li, M. Song, Y. Hao, Z. Yang, J. Zhang, G. Chen, J. Gu, J. Li, X. Qu, and Y. Cheng (2025)

OpenThinkIMG: learning to think with images via visual tool reinforcement learning

.

ArXiv

abs/2505.08617

.

External Links:

Link

Cited by:

§1

,

§2

.

W. Tan, J. Li, J. Ju, Z. Luo, R. Song, and J. Luan (2025)

Think silently, think fast: dynamic latent compression of LLM reasoning chains

.

In

The Thirty-ninth Annual Conference on Neural Information Processing Systems

,

External Links:

Link

Cited by:

§1

.

Q. Team (2025)

Qwen2.5-vl

.

External Links:

Link

Cited by:

§4

.

J. Teoh, M. Tomar, K. Ahn, E. S. Hu, P. Sharma, R. Islam, A. Lamb, and J. Langford (2025)

Next-latent prediction transformers learn compact world models

.

External Links:

2511.05963

,

Link

Cited by:

§3.5

,

§3.5

.

S. Tong, Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie (2024)

Eyes wide shut? exploring the visual shortcomings of multimodal llms

.

In

Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

,

pp. 9568–9578

.

Cited by:

§B.3

,

§4

.

B. Wu, E. Meij, and E. Yilmaz (2025a)

A joint optimization framework for enhancing efficiency of tool utilization in LLM agents

.

In

Findings of the Association for Computational Linguistics: ACL 2025

,

W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.)

,

Vienna, Austria

,

pp. 22361–22373

.

External Links:

Link

,

Document

,

ISBN 979-8-89176-256-5

Cited by:

§1

.

M. Wu, J. Yang, J. Jiang, M. Li, K. Yan, H. Yu, M. Zhang, C. Zhai, and K. Nahrstedt (2025b)

VTool-r1: vlms learn to think with images via reinforcement learning on multimodal tool use

.

External Links:

Link

Cited by:

§1

,

§1

.

P. Wu and S. Xie (2023)

V*: guided visual search as a core mechanism in multimodal llms

.

arXiv preprint arXiv:2312.14135

.

Cited by:

§B.3

,

§4

.

Z. Yang, X. Yu, D. Chen, M. Shen, and C. Gan (2025)

Machine mental imagery: empower multimodal reasoning with latent visual tokens

.

External Links:

2506.17218

,

Link

Cited by:

§1

,

§2

,

§3.1

,

§5.4

.

X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen (2024)

MMMU: a massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

.

In

Proceedings of CVPR

,

Cited by:

§1

.

J. Zheng, J. Shen, Y. Yao, M. Wang, Y. Yang, D. Wang, and T. Liu (2025)

Chain-of-focus prompting: leveraging sequential visual cues to prompt large autoregressive vision models

.

In

The Thirteenth International Conference on Learning Representations

,

External Links:

Link

Cited by:

§2

.

Appendix A

Additional Results and Hyperparameter Settings

A.1

Hyperparameter Settings

We set

λ

\lambda

in Equation (

4

) to

0.2

0.2

and the
number of

[PRED]

tokens to

4

4

across all training settings.

All experiments are conducted on either 4 NVIDIA H200 or 6 NVIDIA H100
GPUs. We train on LVR data for 2,500 steps, on PixelReasoner data for
4 epochs, and on ThinkMorph data for 1 epoch, selecting the best
checkpoint based on validation loss. On H200s, we use a per-device
batch size of 4 with gradient accumulation of 4; on H100s, we reduce
the per-device batch size to 2 to fit within memory. For all runs, LoRA
adapters are configured with rank

r

=

64

r=64

and

α

=

128

\alpha=128

.

At inference, we constrain model outputs to the option letter using a
maximum of 4 tokens, enabling straightforward answer parsing and
ensuring fair, direct comparison with LVR

(Li

et al.

,

2025

)

.

A.2

Embedding Visualisation

Figure

4

visualizes t-SNE projections of the embeddings
learned by

Pearl

and a LoRA SFT baseline, both trained on
ThinkMorph data. For

Pearl

, the two views, i.e.,

⟨

ℐ

0

,

𝒬

⟩

\langle\mathcal{I}_{0},\mathcal{Q}\rangle

and

ℛ

\mathcal{R}

, form
coherent, well-separated clusters that align across tasks (Jigsaw and
Spatial Navigation), indicating that the predictive embedding objective
encourages the model to develop shared, task-discriminative
representations of the input and trajectory. By contrast, the SFT
baseline produces fragmented clusters in which the two views are not
consistently aligned, suggesting that next-token prediction alone does
not induce the same degree of structured latent organisation. This
qualitative difference is consistent with

Pearl

’s quantitative
gains and supports the view that the JEPA objective encourages more
semantically meaningful representations than standard fine-tuning.

A.3

Ablation for Next-Latent Prediction

In this section, we show that using next latent predictions for training the hidden states

h

t

h_{t}

to be the

belief states

for representing views, aids model quality.

Figure 4:

T-SNE visualization of the views

I

0

,

Q

I_{0},Q

and

R

R

across tasks for Qwen2.5-VL-7B-Instruct trained with

Pearl

on the left, compared with simple fine-tuning with next-token prediction on the right.

Model

V

∗

V

D

​

A

∗

{}^{*}_{DA}

V

R

​

P

∗

{}^{*}_{RP}

MMVP

Counting

IQ

Jigsaw

Rel. Ref

Spatial Rel

Single-type, single tool call

Pearl

81.5

86.1

74.5

73.5

68.3

28.2

53.1

39.6

89.5

Pearl

w/o

ℒ

NextLat

\mathcal{L}_{\text{NextLat}}

80.1

83.5

75.0

69.3

65.0

24.0

52.7

42.5

89.5

Pearl

w/o

ℒ

NextLat

,

ℒ

JEPA

\mathcal{L}_{\text{NextLat}},\mathcal{L}_{\text{JEPA}}

79.1

82.6

73.7

65.7

67.5

26.7

45.3

33.6

88.8

Table 3:

Ablation of

ℒ

NextLat

\mathcal{L}_{\text{NextLat}}

on the single-type,
single tool call setting (LVR data, Qwen2.5-VL-7B-Instruct). Removing
the next-latent prediction objective consistently degrades performance,
confirming that encouraging hidden states to act as belief states
improves the trajectory embeddings learned by

ℒ

JEPA

\mathcal{L}_{\text{JEPA}}

.

Bold

denotes best result per column.

Table

1

ablates the contribution of

ℒ

NextLat

\mathcal{L}_{\text{NextLat}}

by comparing

Pearl

against a variant trained without this objective.
Removing

ℒ

NextLat

\mathcal{L}_{\text{NextLat}}

leads to consistent degradation across
most benchmarks, with the most notable drops on V

∗

(80.1 vs. 81.5),
V

D

​

A

∗

{}^{*}_{DA}

(83.5 vs. 86.1), MMVP (69.3 vs. 73.5), and IQ (24.0
vs. 28.2). The only benchmark where the ablated variant is competitive
is V

R

​

P

∗

{}^{*}_{RP}

(75.0 vs. 74.5), suggesting that the next-latent objective
is most beneficial for tasks requiring holistic visual understanding
rather than simple relative positioning.

These results support the theoretical motivation for

ℒ

NextLat

\mathcal{L}_{\text{NextLat}}

:
by encouraging hidden states to converge to belief states (sufficient
summaries of past context), the objective produces more informative
encoding targets for

ℒ

JEPA

\mathcal{L}_{\text{JEPA}}

. Without this regularizer,
the final hidden state of the decoder is a less reliable sequence
representation, which in turn weakens the predictive embedding alignment
signal.

ℒ

NextLat

\mathcal{L}_{\text{NextLat}}

therefore acts as a necessary
complement to

ℒ

JEPA

\mathcal{L}_{\text{JEPA}}

rather than a redundant auxiliary
objective.

Appendix B

Examples and Dataset Statistics

Input Image

𝐱

img

\mathbf{x}_{\mathrm{img}}

Transformer Backbone

⋯

\cdots

<lat>

</lat>

image embeddings

text query

latent reasoning tokens

text answer

Tool Output

𝐳

tool

\mathbf{z}_{\mathrm{tool}}

(depth / segm.)

ℒ

latent

\mathcal{L}_{\mathrm{latent}}

supervise

continuous supervision

ℒ

text

\mathcal{L}_{\mathrm{text}}

discrete supervision

Ground-Truth

Text

𝐲

text

\mathbf{y}_{\mathrm{text}}

Image

Encoder

predicted latent states

predicted text

⋯

\cdots

⋯

\cdots

Legend:

Image Embedding

Text Token

Latent Token (continuous)

Special Delimiter

Figure 5:

Training architecture for latent-augmented multimodal reasoning.

The input image

𝐱

img

\mathbf{x}_{\mathrm{img}}

is encoded into image embeddings and
concatenated with text query tokens. The model autoregressively generates a sequence
of continuous

latent reasoning tokens

(delimited by

<lat>

…

</lat>

), followed by discrete text answer tokens.
Latent tokens are supervised with a continuous regression loss

ℒ

latent

\mathcal{L}_{\mathrm{latent}}

against a visual tool output

𝐳

tool

\mathbf{z}_{\mathrm{tool}}

(e.g. a depth map or segmentation mask),
while text tokens are supervised with a standard cross-entropy loss

ℒ

text

\mathcal{L}_{\mathrm{text}}

against ground-truth text

𝐲

text

\mathbf{y}_{\mathrm{text}}

.

B.1

Reconstruction-based Multimodal Latent Reasoning

Figure

5

illustrates the general training architecture
shared by reconstruction-based latent reasoning methods. The input
image is first passed through an image encoder to produce a sequence
of image embeddings, which are concatenated with text query tokens and
fed into a transformer backbone. The model is then trained to
autoregressively generate a sequence of continuous latent reasoning
tokens, delimited by special

<lat>

…

\ldots

</lat>

markers, before switching to discrete text generation to produce the
final answer. The latent tokens are supervised with a continuous
regression loss

ℒ

latent

\mathcal{L}_{\mathrm{latent}}

against
the output of an external visual tool (e.g., a depth map or
segmentation mask), while the text answer tokens are supervised with a
standard cross-entropy loss

ℒ

text

\mathcal{L}_{\mathrm{text}}

against ground-truth text. A key characteristic of this design is the
autoregressive dependency among latent tokens: each generated latent
token is fed back as input to predict the next, effectively requiring
the model to “imagine” the tool output token-by-token before
transitioning back to discrete generation. This training-inference
asymmetry — where many latent tokens are used during training but
only a small fixed number are decoded at test time — is a central
limitation that

Pearl

is designed to avoid.

B.2

Training Data

The LVR dataset

(Li

et al.

,

2025

)

contains over 450k training samples; however, we find
that loss curves plateau well before exhausting the data, and therefore
train for at most 2,500 steps, reporting results on the best
checkpoint. The ThinkMorph dataset

(Gu

et al.

,

2025

)

comprises 6k samples per subset
across four tool types, yielding 24k training samples in total. Its
reasoning steps are notably more verbose than those in LVR or
PixelReasoner, which contributes to the training–test distribution
mismatch observed for smaller models in Section

5

. We utilize the SFT dataset from PixelReasoner

(Su

et al.

,

2024

)

, which contains

∼

4

​

k

\sim 4k

samples with sequential image transformations over

ℐ

0

\mathcal{I}_{0}

, comprising between 0 and 3 sequential crops of the original image.

Below, we provide examples of training samples from Thinkmorph

(Gu

et al.

,

2025

)

, Viscot

(Shao

et al.

,

2024

)

–which was used by LVR

(Li

et al.

,

2025

)

, and PixelReasoner’s

(Su

et al.

,

2024

)

finetuning dataset.

PixelReasoner Example #13: Visual Question Answering with Tool Use

System

You are a helpful assistant. You may call one or more functions to assist with the
user query. Available tools:

•

crop_image_normalized(bbox_2d, target_image)

— zoom into a
bounding-box region of an image.

•

select_frames(target_frames)

— select frames from a video.

For each function call, return a JSON object inside

<tool_call>...</tool_call>

tags.

User

What does the text on the pink lantern say?

A:

Sakura Festival

B:

Cherry Blossom Viewing

C:

Hanami Celebration

D:

Spring Lanterns

Guidelines: Understand the visual information and the query.
Determine whether to use the provided tools. Reason step by step and put your
final answer inside

{}

.

Assistant

The image shows a serene outdoor scene with cherry blossoms in full bloom.
A person holding an umbrella walks across a bridge.
Pink lanterns with text are strung along the area, likely part of a festival.

I will zoom in to read the text on the pink lantern more clearly.

tool_call

{"name": "crop_image_normalized",

"arguments": {"bbox_2d": [0.0, 0.3, 0.4, 0.6], "target_image": 1}}

Tool Result

Here is the cropped image:

Assistant

I can see the lanterns more clearly now, but the text needs a closer look.

tool_call

{"name": "crop_image_normalized",

"arguments": {"bbox_2d": [0.1, 0.1, 0.2, 0.2], "target_image": 2}}

Tool Result

Here is the cropped image:

Assistant

The close-up clearly shows the text on the pink lantern reads

“Sakura Festival”

.

{

A

}

\{A\}

Viscot (LVR) Example

System

You are a helpful assistant.

User

What are the young men doing with their hands?
Provide a short and direct response.

Assistant

To determine what the young men are doing with their hands, I will focus on the hand region of the individuals visible in the image. I will draw a red bounding box around that region at coordinates

[0.482, 0.396, 0.548, 0.45]

.

Tool Result

Assistant

They are looking at their hands.

ThinkMorph Example: Chart Refocus

System

You are a helpful assistant.

User

What was the total amount of towns, cities, and villages in the US in 2019 with population less than 25000?

Assistant

To answer the question about the total number of towns, cities, and villages in the US in 2019 with a population less than 25,000, I need to identify the relevant segments in the chart that correspond to this population range. The visualization displays the number of incorporated places grouped by population brackets. Specifically, I should look for the bars labeled ’Under 10,000’ and ’10,000 to 24,999’, since both categories fall under the threshold of 25,000 residents. No other category applies, as the next smallest bracket, ’25,000 to 49,999’, exceeds the limit set by the query. By highlighting these two horizontal bars, I can clearly indicate the areas that collectively represent all places with populations below 25,000. This focus directly connects to the population constraint in the question and isolates the relevant data needed to compute the total. Therefore, the manipulation involves drawing a red bounding box around both the ’Under 10,000’ and the ’10,000 to 24,999’ rows in the chart.

Tool Result

Assistant

To determine the total number of towns, cities, and villages in the US in 2019 with a population less than 25,000, I first identify the relevant population categories in the chart. The two categories that fit this criterion are ’Under 10,000’ and ’10,000 to 24,999’, which are highlighted with red boxes in the edited visualization. The numbers associated with these categories are 16,410 for ’Under 10,000’ and 1,572 for ’10,000 to 24,999’. Summing these two values gives the total number of incorporated places with populations below 25,000: 16,410 + 1,572 = 17,982. Therefore, the total amount of towns, cities, and villages in the US in 2019 with population less than 25,000 is 17,982.

ThinkMorph Example: Visual Search

System

You are a helpful assistant.

User

What is the fruit that is on the plate called?

Assistant

To answer the question about the fruit on the plate, begin by visually scanning the image for a plate. The plate is located in the lower right portion of the image, on the tabletop in front of the child. Next, look closely at what is placed on the plate. There is a pale yellow, curved object with a thick peel partially removed, exposing the inner fruit. The shape, color, and partially peeled nature are strong visual cues that this is a fruit. The red bounding box should be drawn around this item on the plate to highlight it as the focal point for identifying the fruit.

Tool Result

Assistant

With the fruit on the plate now highlighted by the red bounding box, it’s easy to closely examine its characteristics. The fruit has a distinctive elongated, curved shape and a pale yellow peel that is partially removed, revealing the soft, creamy flesh inside. These visual details—such as the color, the way the peel splits away from the fruit, and the overall shape—are classic identifiers for a banana. By focusing on the region inside the red bounding box, we can confidently determine that the fruit on the plate is a banana.

ThinkMorph Example: Spatial Navigation

System

You are a maze solver. Your goal is to guide a player from the start to the goal on a grid map while avoiding holes. The player can move one square at a time in the directions left (L), right (R), up (U), or down (D). The frozen lake is not slippery; the player will always move in the intended direction. Moving off the edge or falling into a hole results in failure. Reaching the goal means success.

User

Provide your solution as a sequence of moves wrapped in

boxed

{}, such as {L,R,U,D}. The moves should be comma-separated.

Assistant

To get a clear sense of the challenge, I should first lay out the details of the maze. The grid is 6 rows by 6 columns. The starting point is in the bottom right corner, specifically at row 5, column 5 if counting from the top-left as (0,0). The goal is near the top left, at row 1, column 1.

There are several ice holes scattered around the grid:

•

•

The first row has holes at positions (0,0) and (0,2).

•

The second row contains two holes: one at (1,0) and another at (1,5).

•

The fourth row has a hole at (4,3).

•

The last row (row 5) has a hole at (5,0).

All other squares are safe frozen surfaces. Time to move on to plotting the solution and finding the optimal path.

Tool Result

Assistant

Reviewing my solution, the red path begins at the starting square in the bottom right and follows the moves: Left, Up, Left, Left, Left, Up, Up.

Step by step:

•

First, move left to avoid the edge.

•

Then, up to stay clear of the hole below.

•

Continue left three times, traversing safe frozen squares and bypassing the hole above.

•

Go up twice to reach the goal square near the top left, successfully avoiding all holes and edges.

Each move lands on a safe frozen surface, and the path finishes directly on the goal. This confirms the route is valid and optimal.

Final solution: {L,U,L,L,L,U,U}

B.3

Evaluation Data

We draw our evaluation tasks from MMVP

(Tong

et al.

,

2024

)

, V*

(Wu and Xie,

2023

)

,
and subsets of the Blink dataset

(Fu

et al.

,

2025

)

. Table

4

provides a brief
description of each dataset along with the number of evaluation
samples.

Dataset

Description

No. Samples

V*

Object attributes

191

MMVP

Perception Robustness

300

Counting

Counting objects

120

IQ

Pattern Matching

150

Jigsaw

Multi-image jigsaw resolution

150

Relative Reflectance

Perception

134

Spatial Relation

Relation between objects

143

Table 4:

Overview of datasets and number of evaluation samples.

Appendix C

Limitations and Future Work

While

Pearl

demonstrates strong performance across three
training regimes and multiple model scales, several limitations remain.

Planning over Learned Embeddings.

Pearl

internalizes the

effects

of tool use in the latent
space but does not explicitly plan sequences of actions at inference.
The predictive embeddings learned by

ℒ

JEPA

\mathcal{L}_{\text{JEPA}}

encode
a holistic representation of the full expert trajectory, which implicitly
captures planning structure, but the model does not reason step-by-step
over these representations at test time. A natural extension would be to
use the learned embeddings as a latent world model for explicit multi-step
planning, enabling the model to reason about longer action horizons
without invoking external tools.

Interpretability.

Because

Pearl

operates entirely in a continuous embedding space,
the learned trajectory representations are not directly interpretable.
Unlike reconstruction-based methods, which at least nominally produce
latent tokens aligned with intermediate image edits,

Pearl

makes
no claim about what individual dimensions of the embedding encode. While
the t-SNE visualizations in Figure

4

confirm that
the representations are structured and task-discriminative, understanding

what

the model has internalized about tool use remains an open
question. Developing probing methods or disentangled representations that
make the learned latent structure more transparent is an important
direction for future work.

Training Data Coverage.

Our experiments are limited to settings for which open-source expert
trajectory data exists. In particular, we do not evaluate the
multiple-type, multiple tool call setting due to the absence of suitable
training data. As more diverse trajectory datasets become available, we
expect

Pearl

’s trajectory-level embedding objective to generalize
naturally to richer tool-use settings, given that it places no constraints
on the number or type of tools present in

ℛ

\mathcal{R}

.

Training Cost.

Pearl

requires two separate forward passes per training example
to encode the input view and the trajectory view independently, which
roughly doubles the training-time compute relative to standard SFT.
While this overhead does not affect inference, it may be prohibitive at
very large scale. Exploring more efficient encoding strategies, such as
shared encoders with cross-view masking or cached trajectory embeddings, is a promising avenue for reducing this cost.

BETA