Title: 2604.10545
ArXiv: 2604.10545

Enhanced Self-Learning with Epistemologically-Informed LLM Dialogue

Title:

Content selection saved. Describe the issue below:

Description:

License: CC BY 4.0

arXiv:2604.10545v1 [cs.HC] 12 Apr 2026

Enhanced Self-Learning with Epistemologically-Informed LLM Dialogue

Yi-Fan Cao

0000-0002-5892-5052

Hong Kong University of Science and Technology

Hong Kong

China

ycaoaw@connect.ust.hk

,

Kento Shigyo

0000-0002-5095-7500

Hong Kong University of Science and Technology

Hong Kong

China

kshigyo@connect.ust.hk

,

Yitong Gu

0009-0001-2890-5448

Hong Kong Baptist University

Hong Kong

China

yitonggu@life.hkbu.edu.hk

,

Xiyuan Wang

0009-0008-1839-2010

ShanghaiTech University

Shanghai

China

wangxy7@shanghaitech.edu.cn

,

Weijia Liu

0009-0001-6332-5705

Hong Kong University of Science and Technology (Guangzhou)

Guangzhou

China

wliu383@connect.hkust-gz.edu.cn

,

Yang Wang

The University of Hong Kong

Hong Kong

China

yang.wang@hku.hk

0000-0002-8903-2388

,

David Gotz

0000-0002-6424-7374

University of North Carolina at Chapel Hill

Chapel Hill

North Carolina

USA

gotz@unc.edu

,

Zhilan Zhou

0000-0003-1236-1287

University of North Carolina at Chapel Hill

Chapel Hill

North Carolina

USA

zzl@cs.unc.edu

and

Huamin Qu

Hong Kong University of Science and Technology

Hong Kong

China

huamin@ust.hk

0000-0002-3344-9694

Abstract.

Large Language Models (LLMs) have advanced self-learning tools, enabling more personalized interactions. However, learners struggle to engage in meaningful dialogue and process complex information. To alleviate this, we incorporate epistemological frameworks within an LLM-based approach to self-learning, reducing the cognitive load on learners and fostering deeper engagement and holistic understanding.

Through a formative study (N=26), we identified epistemological differences in self-learner interaction patterns. Building upon these findings, we present

CausaDisco

, a dialogue-based interactive system that integrates Aristotle’s

Four Causes

framework into LLM prompts to enhance cognitive support for self-learning.

This approach guides learners’ self-learning journeys by automatically generating coherent and contextually appropriate follow-up questions. A controlled study (N=36) demonstrated that, compared to baseline,

CausaDisco

fostered more engaging interactions, inspired sophisticated exploration, and facilitated multifaceted perspectives. This research contributes to HCI by expanding the understanding of LLMs as educational agents and providing design implications for this emerging class of tools.

Large Language Models (LLMs), Sensemaking, Human-AI Interaction, Epistemology, Self-Learning

†

†

copyright:

acmlicensed

†

†

ccs:

Human-centered computing Interactive systems and tools

†

†

ccs:

Applied computing Education

Figure 1.

CausaDisco

provides users with original learning materials (A) and a core concept graph (B) to facilitate self-learning.
Once a user initiates a dialogue, the LLM chatbot offers preliminary answers and then automatically generates four epistemologically-informed follow-up questions (C) to encourage deeper exploration. Concurrently,

CausaDisco

creates a query tree map (D) to assist users in managing their conversation records.
When users are satisfied with the results of multi-turn dialogues (T1-3), they can proceed to explore and learn about new topics.

This figure presents the interface and interaction design of CausaDisco, which enhances self-learning by providing original materials and a core concept graph for guidance. When users start a dialogue, the LLM chatbot offers initial answers and generates four follow-up questions to promote deeper understanding. Simultaneously, a query tree map is created to help users manage conversation records. Once users are satisfied with multi-turn dialogues, they can explore new topics.

1.

Introduction

The rise of Large Language Models (LLMs) has revolutionized education, introducing a new generation of dialogue-driven learning tools

(Luo

et al.

,

2022

; Chen

et al.

,

2023b

; Zhang

et al.

,

2023

)

.
These tools have greatly improved the efficiency of independent learning, supporting a wide range of tasks, from spoken language practice

(Huang

et al.

,

2022

)

and text editing

(Kim,

2023

)

to literature searches

(Zheng

et al.

,

2024a

,

b

)

and information summarization

(Reddy and Guha,

2023

)

.
With the adoption of prompt engineering techniques, the potential of LLM-based educational tools to facilitate deeper learning and knowledge construction has been further explored

(Suh

et al.

,

2023

)

.
These dialogue-driven tools can help streamline the cognitive processes involved in sensemaking, enabling learners to acquire, comprehend, integrate, and apply new knowledge more effectively

(Abdelghani

et al.

,

2024

; Gero

et al.

,

2024

)

.

Although LLM-based educational tools promote a flexible and intuitive “search-as-learning” sensemaking process

(Liu

et al.

,

2024a

)

, many still require substantial effort from learners to engage in meaningful dialogue and process complex information.
This is particularly challenging when using less customized LLM tools, as learners often spend considerable time crafting precise questions to elicit responses aligned with their specific needs.
Such exploratory and iterative dialogue-driven interactions can be especially demanding for less proactive learners accustomed to more passive input-driven sensemaking.
Consequently, even with LLM-based tools, many learners still encounter substantial, yet often hidden, cognitive load.

Extensive studies have focused on enhancing the contextuality

(Akhtar

et al.

,

2019

; Shihab

et al.

,

2023

; Chen

et al.

,

2023a

)

and workflow

(Suh

et al.

,

2023

; Gao

et al.

,

2024

; Shankar

et al.

,

2024

)

of LLM-based educational tools to support learners’ sensemaking processes.
However,

existing work often centers

on domain-specific applications

(Sheng

et al.

,

2023

; Chen

et al.

,

2024b

)

or caters to early adopters already proficient in strategically engaging with LLM chatbots

(Zheng

et al.

,

2024b

)

.

This focus overlooks a fundamental aspect of human-AI interaction: sensemaking is a highly individualized cognitive process.
This process is shaped by one’s

epistemological schema

, which represents structured beliefs about the

nature

,

source

, and

justification

of knowledge

(Hofer and Pintrich,

1997

; Brownlee

et al.

,

2002

; Bråten,

2010

; Baker and Anderman,

2020

)

.

Consequently, a critical gap exists in understanding how diverse mental models, reflected in different LLM interaction behaviors influence the design of effective, usable, and inclusive educational tools.

To bridge this knowledge gap, we conducted a formative study with 26 participants.

Data collection was triangulated

(Carter,

2014

)

through dialogue records captured by a probe system, online surveys, and semi-structured interviews

1

1

1

Please refer to the supplementary materials for details of our survey and semi-structured interview protocols.

. Through a thematic analysis of the collected data, we derived a query taxonomy that revealed distinct interaction behaviors, while also examining participants’ underlying epistemological schemas. The synthesis of these analyses identified four

interaction patterns

:

proactive

,

validation-seeking

,

content-focused

, and

receptive

.
We revealed key differences by analyzing these patterns within the epistemological framework of knowledge nature, source, and justification, particularly between proactive and receptive patterns.
The proactive pattern was characterized by active knowledge construction, evidenced by participants formulating effective questions, requesting elaborations, and synthesizing information. Conversely, the receptive pattern was characterized by fewer questions, less exploration, and a preference for receiving direct answers from LLMs rather than actively building understanding. Furthermore, the proactive pattern was associated with high-performing participants, while the receptive pattern was often observed in those who struggled in the learning process.

The observed epistemological differences and learning outcomes across different patterns highlight the need for tools to accommodate diverse learner needs. We identified three primary challenges in designing inclusive learner-LLM interactions

:

limited interactivity

,

inefficient information verification

, and

confirmation bias

.
Addressing these challenges requires educational tools that support: (1) lower cognitive load, (2) sustained engagement, (3) multidimensional thinking, (4) metacognition, and (5) validation assistance.
Guided by these requirements, we developed

CausaDisco

, a proof-of-concept system designed to support more effective learner-LLM interactions.

CausaDisco

comprises four main views (see Fig.

1

): an

Embedded Content View

providing access to original learning materials; a

Concept Graph

visualizing relationships between core concepts; a

Q&A Conversation View

prompting epistemologically-informed follow-up questions; and a

Tree Map View

displaying the underlying query logic.

Specifically, in the

Q&A Conversation View

, we incorporate epistemological frameworks with LLM prompt engineering to enhance self-learners sensemaking.

This approach is inspired by pedagogical research demonstrating the benefits of structured, epistemologically-informed frameworks for critical thinking, analytical depth, and knowledge acquisition in self-learning contexts

(Peters,

2000

; Bråten,

2010

; Krasmann,

2020

)

.
Drawing on these insights, as well as

the interaction patterns distilled from

our formative study, we identified the need for a systematic approach to generating follow-up questions that support comprehensive sensemaking.
We selected Aristotle’s

Four Causes

(Hocutt,

1974

; Falcon,

2006

)

–a classical epistemological framework for knowledge construction–due to its remarkable alignment with our empirically derived query taxonomy.
This framework prompts users to consider the

Material

,

Formal

,

Efficient

, and

Final

causes of concepts, effectively mapping to the diverse types of questions learners naturally ask.

To evaluate the efficacy of

CausaDisco

, we conducted a controlled user study with 36 participants employing a between-subjects design.
We assessed both objective and subjective measures, including quiz scores and Likert-scale ratings to assess engagement, efficiency, comprehensiveness, sophistication, and usability of the system.
Our results indicate that

CausaDisco

significantly facilitates users’ sensemaking process, fostering a more engaging learning experience compared to the baseline condition.
In summary, our core contributions are as follows:

•

A taxonomy of interaction patterns in self-learning with LLMs, identifying key epistemological differences and outlining design requirements for supporting user sensemaking.

•

CausaDisco

, an interactive proof-of-concept system for self-learning. The system leverages Aristotle’s

Four Causes

theoretical framework for LLM prompt engineering to guide user sensemaking and includes features for knowledge synthesis and learning journey tracking.

•

Empirically-grounded design implications for dialogue-driven LLM educational tools, derived from a within-subjects evaluation of

CausaDisco

’s efficacy in enhancing interactivity and sensemaking.

2.

Related Work

Given the interdisciplinary nature of this research, our literature review integrates technical, epistemological, and pedagogical perspectives. It is structured as follows:

1)

LLM-based educational agents

; 2)

Integrating epistemology for LLM sensemaking

; and 3)

Toward effective learner-LLM interactions

.

2.1.

LLM-Based Educational Agents

LLM-based educational agents advance technology-enhanced learning (TEL) by creating personalized learning experiences

(Gan

et al.

,

2023

; Lin,

2023

; Huber

et al.

,

2024

)

.

These agents improve intelligent tutoring systems (ITS) through real-time assistance

(Li

et al.

,

2024a

)

, customized feedback

(Fu

et al.

,

2024

)

, and adaptive learning paths

(Lee

et al.

,

2024

)

. Such features empower learners to regulate their cognitive processes and support various tasks, including self-reflection

(Li

et al.

,

2023

)

, goal-setting

(Deng

et al.

,

2023

)

, and problem-solving

(Kumar

et al.

,

2023

)

.

This makes LLM-based educational agents particularly valuable in self-learning contexts, where individualized and adaptive support is critical.

Self-directed learning is a process involving continuous cognitive and behavioral adjustment aimed at understanding specific knowledge domains

(Demirbag,

2021

; Huang

et al.

,

2023

)

. Learning outcomes can vary depending on individual strategies, and pedagogical research emphasizes the importance of active engagement and self-regulation in this process

(Wang

et al.

,

2015

; Wekerle

et al.

,

2022

)

. Recent HCI studies have explored how LLM-based educational agents can enhance the effectiveness of self-learning by promoting autonomy, particularly from students’ perspectives

(Limna

et al.

,

2023

; Li

et al.

,

2024b

)

.
For instance, Goslen

et al.

(

Goslen

et al.

,

2024

)

proposed an LLM-based system that automatically generates student planning strategies based on historical interaction data, enabling dynamic support in game-based learning environments.
Through fine-tuning for domain-specific instruction, these agents have demonstrated success in areas such as language learning

(Liu

et al.

,

2024c

)

, technology literacy

(Gao

et al.

,

2024

; Hartley

et al.

,

2024

)

, and research ideation

(Liu

et al.

,

2024b

)

. These advances enable self-learners to effectively manage their cognitive processes and improve practical skills with reduced dependence on human guidance.

Despite these advances, LLM-based educational agents have yet to achieve the dynamic reciprocal interaction where learners and systems mutually shape each other’s development, as envisioned by pedagogical scholars

(Schulte and Budde,

2018

)

.
This limitation stems primarily from the agents’ inability to guide users in crafting effective prompts

(Shah and Bender,

2022

; Fiannaca

et al.

,

2023

; Zamfirescu-Pereira

et al.

,

2023

)

, which increases cognitive load and hinders knowledge acquisition.
Although features such as automatic question generation

(Zheng

et al.

,

2024b

; Hu

et al.

,

2024

)

and dialogue history management

(Safranek

et al.

,

2023

)

improve interaction intuitiveness, a lack of systematic guidance limits the ability of self-learners—especially those unfamiliar with AI or interactive learning—to fully harness these agents’ potential.

Our study addresses this gap by investigating how different mental models shape self-learners interaction patterns with LLM-based educational agents. These insights will inform the design of more effective, accessible, and inclusive AI-enhanced educational systems, ultimately enhancing learner engagement for a broader audience.

2.2.

Integrating Epistemology for LLM Sensemaking

The advent of LLMs has spurred their integration into tools augmenting human sensemaking during self-learning.

Sensemaking is a vital cognitive process that influences how learners process information, ultimately impacting their learning outcomes and efficiency

(

Marchionini

,

2019

; Rong

et al.

,

2023

)

.

This cyclical process encompasses several cognitive stages:

acquiring

,

comprehending

,

integrating

, and

applying

new knowledge

(Pirolli and Card,

2005

)

. Recognizing LLMs’ potential to transform this process, the HCI research community has explored various approaches to enhance the different cognitive stages

(Lee and Ma,

2024

; Gero

et al.

,

2024

; Ma

et al.

,

2024

)

.
For instance, Zheng

et al.

(Zheng

et al.

,

2024b

)

enhanced knowledge search and acquisition through an interdisciplinary literature search system that facilitates cross-domain synthesis. Building on this foundation, Suh

et al.

(Suh

et al.

,

2023

)

advanced knowledge comprehension by developing an interactive system that reveals multilevel abstractions of complex information spaces. Taking a different approach, Jin

et al.

(Jin

et al.

,

2024

)

focused on knowledge construction by leveraging LLMs as teachable agents within a “learning by teaching” paradigm.

Despite their promise, existing LLM sensemaking applications rarely consider how individual epistemological schemas shape knowledge construction.
Epistemology–the philosophical study of knowledge and belief that guides learners to evaluate the credibility and relevance of information sources–underpins how learners make sense of new information

(

Hofer and Pintrich

,

1997

; Schommer-Aikins

,

2004

; Shirzad

et al.

,

2022

)

.
Pedagogical research emphasizes that the effectiveness of sensemaking depends fundamentally on learners’ epistemological schemas–their structured beliefs about knowledge’s

nature

,

source

, and

justification

(

Hofer and Pintrich

,

1997

; Bråten

,

2010

; Baker and Anderman

,

2020

)

. These beliefs critically shape how individuals process information during learning, from initial interpretation to final evaluation

(

Zhang and Soergel

,

2020

)

. For example, Baker

et al.

(

Baker and Anderman

,

2020

)

demonstrate that such personal epistemological schemas influence the entire spectrum of self-learning, affecting learners’ motivation, achievement levels, and engagement with learning materials.

Given the fundamental role of epistemology in knowledge construction, prior research has investigated applying epistemological frameworks to AI models for educational tools

(

Woolf

et al.

,

2013

; Zhang

et al.

,

2023

)

.
For instance, Rosé

et al.

(

Rosé

et al.

,

2008

)

proposed an AI-driven framework grounded in an understanding of epistemology, which can trace learners’ knowledge construction processes, thereby improving learning outcomes.
However, the rapid adoption of LLMs has shifted recent scholarly attention toward their capacity to disrupt traditional teaching practices and introduce epistemic paradoxes

(

Sibilin

,

2023

; Chavanayarn

,

2023

; Cassinadri

,

2024

)

. Consequently, the potential of leveraging LLMs for accommodating learners’ epistemological schemas, which is crucial for sensemaking, remains largely unexplored.

Our research addresses this gap by integrating epistemological frameworks into LLM-based systems through prompt engineering, guiding users through the processes of problem framing, evidence evaluation, and reflection, key stages of sensemaking.

This domain-agnostic approach ensures versatility and offers a structured yet adaptable framework for knowledge exploration.

2.3.

Toward Effective Learner-LLM Interactions

Understanding and improving the interactivity in learning activities remains a persistent focus in pedagogy.
The

Interactive

,

Constructive

,

Active

, and

Passive

(ICAP) framework, proposed by Chi

(Chi,

2009

)

, categorizes learner engagement modes based on observable behaviors and underlying cognitive processes, providing a foundation for understanding learner interactivity.
A subsequent five-year study

(Chi and Wylie,

2014

)

demonstrated the relative effectiveness of these engagement modes, ranking them from most to least effective: interactive, constructive, active, and passive.
This finding has fueled further learner-centered research exploring the relationship between engagement modes and learning outcomes in specific educational scenarios

(Wang

et al.

,

2015

; Schulte and Budde,

2018

)

.
While some variations exist in the resulting classifications and analyses of contributing factors, the principle that “

good learning is active learning

” remains widely accepted

(Brown

et al.

,

2014

)

.
However, translating the ICAP framework into practice, particularly in designing interactive and constructive activities, has proven challenging

(Chi

et al.

,

2018

)

.

Advances in technology create new opportunities to apply the ICAP framework.
Recent educational research has extended the ICAP theory to the context of TEL

(Sailer

et al.

,

2021

; Wekerle

et al.

,

2022

; Sailer

et al.

,

2024

)

. Building on the ICAP framework, Sailer

et al.

(Sailer

et al.

,

2024

)

introduced the

Substitution

,

Augmentation

,

Modification

, and

Redefinition

(SAMR) model to classify levels of technology integration in learning. The introduction of LLM-based educational agents marks a transformative step toward SAMR’s “redefinition” level, especially in self-learning scenarios. These agents enable scalable conversational interactions that can enhance learner engagement.

Although the conversational interface of LLM-based educational agents suggests simplicity and intuitive interaction similar to human communication

(Smestad and Volden,

2019

)

, extant HCI literature reveals notable disparities in how learners benefit from these tools

(Chen

et al.

,

2024a

)

.
This variability often arises from learners’ diverse motivations, levels of LLM literacy, mental models, and educational backgrounds, ultimately influencing their interaction strategies and learning outcomes

(Mogavi

et al.

,

2024

; Zamfirescu-Pereira

et al.

,

2023

)

.

For instance, research on help-seeking behaviors among self-learners highlights how self-efficacy, perceived task difficulty, and tool accessibility influence their resource utilization

(Moores and Chang,

2009

; Karabenick and Dembo,

2011

; Glassman

et al.

,

2015

; Urgo and Arguello,

2022

)

. Overconfident learners, in particular, may only seek external support when facing substantial challenges or when perceiving the tool as beneficial for their understanding

(Moores and Chang,

2009

)

.

This tendency can hinder their ability to accurately assess their learning progress. Furthermore, novice AI users often struggle to craft effective prompts during LLM interactions, compounding the challenges they face

(Zamfirescu-Pereira

et al.

,

2023

; Fiannaca

et al.

,

2023

)

.

To promote effective learner-LLM interaction, designing more inclusive LLM-based tools that provide effective cognitive support and accommodate the diverse needs of self-learners is essential

(Cohn

et al.

,

2024

)

.
Drawing on the ICAP framework

(Chi and Wylie,

2014

)

and our previously discussed epistemological perspective, our research examines how self-learners interact with LLM-based educational agents in practice. By analyzing these interaction patterns and their underlying mental models, we explore how to enhance LLM prompt engineering to better facilitate learner interactivity and knowledge co-construction.

3.

Formative Study

This section describes our formative study, which adopts a user-centered approach

(Vredenburg

et al.

,

2002

)

. The study investigates participants’

interaction behaviors

,

epistemological schemas

, and the

challenges

they face during self-directed learning. By synthesizing these dimensions, we derive a categorization of participants’

interaction patterns

and deduce the latent

mental models

that guide their engagement when learning new topics. These findings inform the design priorities of our system.

Specifically, we outline the recruitment process and participant demographics, describe the study procedure, and explain the data analysis methods.
This study was conducted in accordance with all relevant ethical guidelines and received full approval from our university’s Institutional Review Board (IRB). Prior to data collection, verbal informed consent was obtained from all participants.

3.1.

Recruitment and Participants

Recruitment

We employed a dual-pronged approach, combining convenience and snowball sampling strategies, to gather a representative sample of target users who use LLM-based chatbots for self-learning.
We disseminated recruitment messages with contact information through multiple channels, including social media, word of mouth, and campus bulletin boards at five institutions of higher education in East Asia and North America.
Our recruitment criteria consisted of the following: (1) adults aged 18 years or older, (2) individuals capable of interacting with LLM-based chatbots, (3) those with at least three months of prior experience using LLMs for self-learning, and (4) participants without prior knowledge of blockchain technology, which served as the primary reading material for our self-learning task.
We received a total of 35 applications and conducted a thorough screening of respondents based on our criteria.
We stopped the recruitment process when data analysis results reached theoretical saturation

(Clarke and Braun,

2017

)

, meaning no new insights emerged from participants’ dialogue records or semi-structured interviews (SSI).
Ultimately, we recruited 26 participants with varied educational backgrounds, from various disciplines, and from seven countries and regions (see Table

6

in Appendix

A

).

Table 1.

Descriptive statistics of participants in the formative study. This diverse sampling pool provides insights into design requirements for improving LLM-based chatbots for self-learning across various demographics.

Participants

The 26 participants were between 23 and 32 years of age (M = 27.04, SD = 2.28), and carefully selected to represent diverse backgrounds and experiences with LLM-based tools (see Table

1

).
Sixteen held a master’s degree, followed by four with bachelor’s degrees and six with doctorates, indicating a strong representation of individuals engaged in advanced studies or research.
Our participants represented a diverse range of disciplines, with a substantial portion (13) coming from STEM fields–a group shown to have positive perceptions of LLMs as learning tools

(Bernabei

et al.

,

2023

)

.
The remaining participants were from social sciences (4), interdisciplinary studies (5), and other fields (4), ensuring broader perspectives.
Most participants (22) had experience in academic training, providing a foundation for evaluating LLM-based self-learning tools.
While ten participants had 1-2 years of experience using LLM tools, a significant portion (15) had been using them for less than a year, and only one for over two years, indicating a mix of early adopters and more recent users.
Finally, participants’ familiarity with LLM-based chatbots varied: 21 reported being somewhat (11), very (6), or extremely (4) familiar, and five were unfamiliar.
This distribution enabled comparison between more and less adept chatbot users, a crucial aspect of our study design.

3.2.

Formative Study Design

The formative study aimed to investigate

participants’

interaction behaviors

,

epistemological schemas

, and the

challenges

they encountered

during self-learning through a task-focused probe system.
Before formally launching the formative study, a pilot study was conducted with two team members from diverse disciplinary backgrounds to validate the process and estimate the required duration.

3.2.1.

Study Setup

The formative study allowed both in-person (n=12) and online (n=14) participation.
In-person sessions were held at the recruiting institutions. Online sessions were conducted via Zoom

2

2

2

https://zoom.us/

to facilitate real-time interaction and observation of participant behavior.
To provide consistent participant experiences within a controlled learning environment, we employed Qualtrics

3

3

3

https://www.qualtrics.com/

for online surveys and a task-focused probe system

accessed via the Web directly

.
This system included raw learning materials, a textual article, and an LLM-chatbot limited to responding only to the provided learning material (see Fig.

10

in Appendix

A.2

).
This controlled setup ensured a fair self-learning process, free from misinformation and external online sources.

Aiming for smooth operation, participants received essential documents, including the consent form and study guidance, one day prior to their scheduled session.
These documents detailed the research goals and experimental procedures and addressed any potential questions.
Notably, we did not share the probe system link in advance to avoid pre-exposure to the learning materials, which could potentially bias our assessment of user learning performance.
The formative study sessions were allocated 90 minutes based on the results of the pilot study.

Figure 2.

Formative Study Workflow (N=26): A three-phase investigation of LLM interaction patterns during self-learning. (A)

Data Collection:

Gathered dialogue records, quiz scores, survey responses, and transcripts of semi-structured interviews. (B)

Thematic Analysis:

Classified four distinct LLM interaction patterns (proactive, validation-seeking, content-focused, receptive). (C)

Epistemological Analysis and Design Implications:

Compared and integrated epistemological frameworks across the identified interaction patterns, revealing insights into participant challenges and informing the design requirements for an epistemologically-grounded LLM-assisted learning system.

3.2.2.

Study Procedure

Upon receiving informed consent, the study proceeded with the following steps (Fig.

2

A):

Self-Learning Task

Participants initially completed a sensemaking task using an informational article on non-fungible tokens (NFTs)

(Creighton,

2023

)

.
Using a think-aloud protocol, participants provided real-time feedback on the probe system’s design and usability during their interactions.
Participants self-regulated their learning pace, ceasing interaction once they felt they had adequately understood the article.
They then downloaded a record of their dialogue with the LLM for future reference.
Task completion time ranged from 15 to 35 minutes (M = 20.21, SD = 5.12).

Quiz

Participants then completed a 15-minute, ten-point quiz assessing the comprehensiveness and accuracy of their acquired knowledge. The quiz, consisting of four multiple-choice and two short-answer questions, was designed by two NFT domain experts who also independently graded the responses.
The final scores were the average of the two experts’ grades.

Survey and Semi-Structured Interview

To further understand individual learning experiences, a seven-point Likert scale survey was administered, gathering data on LLM usage patterns (e.g., familiarity, modalities) and perceived challenges.
Semi-structured interviews then explored individual learning processes in depth, with participants sharing the rationale behind their query chains, challenges, and suggestions for improving LLM interactions.

4

4

4

For illustrative examples of conversation records corresponding to each interaction pattern, please refer to the supplementary materials.

Each interview lasted between thirty minutes and one hour, with audio recordings and verbatim transcriptions conducted.

Figure 3.

This figure presents our six-step process for analyzing participants’ interaction behaviors and epistemological frameworks based on data collected from our formative study.

3.3.

Data Analysis

We employed reflexive thematic analysis

(Braun and Clarke,

2019

,

2021

)

to explore participants’ sensemaking processes and categorize their interaction patterns.
Building on existing literature

(Aliannejadi

et al.

,

2021

; Schneider

et al.

,

2023

)

, we defined interaction patterns based on two key aspects. The first focuses on participants’

interaction behaviors

, as reflected in the queries they generated during the self-learning process.
 The second dimension pertains to their

epistemological schemas

, derived from their chain of queries and self-reported knowledge construction processes detailed in the interview transcripts.
We also analyzed the

challenges

participants encountered while interacting with the LLM through a detailed examination of the interview transcripts, which was further complemented by their survey responses. The complete process and use of collected data sources are illustrated in Fig.

2

B.

Given the exploratory nature of our study within the emerging field of LLM-mediated self-learning, we adopted a predominantly

inductive

approach.
This approach enabled us to develop a query taxonomy of learner-LLM interactions while drawing on an epistemological framework

(Aksan,

2009

; Mokhtari,

2014

; Huang

et al.

,

2023

)

to analyze participants’ knowledge construction processes.
This flexible approach enabled us to explore the multifaceted nature of user experiences in depth

(Guest

et al.

,

2011

)

, while grounding our analysis in relevant theoretical constructs.

Three analysts participated: two authors (who observed all self-learning tasks and conducted all interviews) and one NFT domain expert.

3.3.1.

Interaction Behaviors Analysis

Initially, the three analysts independently examined participants’ interaction behaviors within the self-learning task (see Fig.

3

A).

The goal was to systematically categorize 179 participant-generated queries based on the

knowledge dimensions

explored and the

conversation strategies

employed. The analysts met regularly to share coding results, discuss discrepancies, and refine the codebook until reaching a consensus for the query taxonomy.
The resulting query taxonomy not only highlights the types of queries generated by participants but also captures the frequency of their interactions and the diversity of conversation strategies they employed.

Based on this analysis, participants’ interaction behaviors were mapped along a
spectrum ranging

from

prompt-naive

(low interaction frequency; few exploration dimensions and conversation strategies) to

AI-interactive

(high interaction frequency; diverse exploration dimensions and conversation strategies) (see Fig.

3

B).
To further contextualize these patterns, the analysts conducted a detailed examination of each participant’s chain of queries (see Fig.

3

C). This examination informed the subsequent semi-structured interviews designed to elicit a nuanced understanding of participants’ epistemological schemas for self-learning.

3.3.2.

Epistemological Schemas Analysis

To investigate participants’ epistemological schemas, we integrated the previously established query taxonomy, the analysis of query chains, and the data gathered from the semi-structured interviews. During these interviews, we asked participants about their information needs and how they constructed meaning, referencing their specific query flows from the dialogue records.
We analyzed the interview data (22.5 hours recorded, transcribed into 26 transcripts) alongside the corresponding dialogue records.
This analysis employed a three-dimensional framework focusing on the nature, source, and justification of knowledge to understand participants’ personal epistemological schemas

(Aksan,

2009

; Mokhtari,

2014

; Huang

et al.

,

2023

)

.
Guided by this framework, we systematically coded the interview data within

Atlas.ti

5

5

5

https://atlasti.com/

software (see Fig.

3

D).
We met weekly to review findings, resolve disagreements, and refine the codes until a consensus was reached.
As a result, we placed participants’ epistemological schemas along a spectrum from

reflective-minded

to

confirmatory-oriented

(see Fig.

3

E). Reflective-minded patterns are characterized by a dynamic view of knowledge, knowledge construction through exploration, and justifications through multi-faceted validation. In contrast, confirmatory-oriented patterns demonstrate an authoritative view of knowledge, reliance on external expertise, and justification by seeking standard answers.

3.3.3.

Interaction Patterns Classification

This mixed-methods approach revealed two key spectra for categorizing participants based on a two-dimensional analytical framework:

interaction behavior

, ranging from prompt-naive to AI-interactive, and

epistemological schema

, spanning from reflective-minded to confirmatory-oriented. The integration of these dimensions provided a robust and holistic analysis, generating a quadratic classification of interaction patterns (similar to

(He

et al.

,

2023

)

):

proactive

,

validation-seeking

,

content-focused

, and

receptive

(see Fig.

3

F).

Figure 4.

This 2x2 matrix presents four distinct interaction patterns, categorized by AI interactivity and reflective-mindedness: Prompt-Naive, Confirmatory-Oriented, Reflective-Minded, and AI-Interactive. This classification reflects the mental models associated with different learning strategies observed during self-learning.

4.

Findings

This section

unpacks four interaction patterns and their epistemological differences.

These analyses, combined with survey responses and interview data, revealed key challenges and informed the design requirements for our LLM-supported self-learning prototype (see Fig.

2

C).

4.1.

Learner-LLM Interaction Patterns

This section illustrates four learner-LLM interaction patterns.

These patterns are characterized by their query types, conversational strategies, self-reported learning workflows (gathered through semi-structured interviews), and quiz performance.

4.1.1.

Proactive (n = 8)

Eight participants exhibited the proactive interaction pattern marked by high interactivity and reflectiveness and actively pursued knowledge

(see Fig.

4

A).
They viewed the chatbot as a collaborative partner, engaged in extended, multi-round dialogues, and critically examined responses.

The proactive interaction pattern is demonstrated through the participants’ elevated query volume and diverse exploratory query types (see Table

2

).

They delved deeper into topics by asking about “attributes of concepts,” “co-existent concepts,” “realistic application,” and “real-world consequences”–all highlighting a desire to understand a topic from multiple angles.
For example, P9 explained, “

I don’t just want the ‘what,’ I want the ‘why’ and the ‘how does this connect to other things?’

”

In addition, these participants strategically employed queries such as “change perspectives” and “rephrase requests,” actively guiding the conversation toward providing a more holistic understanding.
For instance, P12 revealed, “

I rephrase and ask about the same concept multiple times in different ways to get a more comprehensive and specific response, instead of an overly general one.

”
This demonstrates an active role in shaping the learning experience.
This interaction pattern was often reflected in their high quiz scores, ranging from 7 to 9 out of 10 (M = 8.28, SD = 0.62), and suggested active learning strategies driven by curiosity and a desire for deeper comprehension.

Table 2.

This table presents the query taxonomy developed during the formative study, structured around three key aspects:

Knowledge Dimension

,

Conversation Strategy

, and

Interaction Frequency

(visualized numerically via a heat map).

Questions marked with an asterisk (*) represent follow-up queries designed to provide deeper insights. While the heat map directly quantifies the AI-interactivity dimension, reflective-mindedness—captured through triangulated interview analysis—is only partially represented here.

4.1.2.

Validation-Seeking (n = 6)

Six participants displayed the validation-seeking interaction pattern, characterized by high interactivity but low reflectiveness, primarily engaging with chatbots to confirm existing knowledge

(see Fig.

4

B).

Their high-frequency interaction behaviors primarily involve asking specific questions with predetermined expectations to validate assumptions or test hypotheses (see “validate hypothesis” in Table

2

)

Their repeated queries about “co-existent concepts” and “realistic applications” reflected this desire to align new information with established understanding and real-world scenarios (see Table

2

).
As P1 noted, “

Relating new concepts to hypotheses formed from my existing knowledge boosts my confidence in self-learning.

”

Furthermore, they often apply conversational strategies such as “validate hypothesis” and “assess accuracy,” emphasizing their preference for reinforcing existing knowledge rather than exploring new perspectives.
As P22 shared, “

Instead of constantly trying to synthesize vast new concepts, presenting a hypothesis and seeking clear feedback is more straightforward.

”
Their quiz scores, ranging from 5 to 7.5 out of 10 (M = 6.33, SD = 0.98), suggest that while these participants prioritize accuracy, relying on their existing understanding may not always facilitate the necessary expansion of their knowledge base for a thorough grasp of the subject.

4.1.3.

Content-Focused (n = 6)

Six participants demonstrated the content-focused interaction pattern, characterized by low interactivity but high reflectiveness in their behaviors, prioritizing thorough understanding over fragmented knowledge

(see Fig.

4

D).
These participants often engaged in extensive exploration of learning materials before directly engaging with the LLM chatbot. As P14 articulated, “

I do not want to ask superficial questions.
I need time to process the information and figure out valuable questions

.”
With this pattern, participants tended to use the LLM less often, primarily for retrieving information directly related to specific aspects of the learning materials, such as “components of concepts,” “co-existent concepts,” and “realistic applications.”

This pattern is further evidenced by their sole use of the “summarize content” strategy (see Table

2

).
P17 aptly captured this tension, noting, “

Simultaneously grasping new concepts and forming insightful questions is challenging, so I mainly use the LLM to facilitate understanding

.”
However, this approach may not always translate to high performance within a time-limited assessment context, as suggested by the significant individual variability in quiz scores, ranging from 5 to 8 out of 10 (M = 6.58, SD = 1.64).

4.1.4.

Receptive (n = 6)

The receptive interaction pattern was evident in six participants, whose behaviors were marked by minimal engagement, limited reflectiveness, and passive interaction with the LLM chatbot

(see Fig.

4

C).
Accustomed to traditional learning methods like reading, online courses, and structured video lectures–which typically involve limited interaction–these participants often struggled to formulate precise questions or to fully leverage the chatbot’s capabilities.
As P13 admitted, “

I am not used to learning with an LLM. I do not think asking questions is necessary.

” This sentiment is further evidenced by a lack of discernible patterns in their query topic and strategy preferences, indicating a less systematic LLM-interaction approach (see Table

2

).
Furthermore, their quiz scores, ranging from 5 to 7 out of 10 (M = 5.75, SD = 0.82), generally suggest that those who struggle to adapt to this interactive learning modality may experience less favorable learning outcomes.

4.2.

User-Centered Design

This section illuminates the epistemological differences across the four learner-LLM interaction patterns. Integrating insights from the surveys and interviews,

this analysis reveals the primary challenges (

C1-3

) participants face when using LLMs and informs design requirements (

DR1-5

) for enhancing LLM-supported self-learning tools.

Table 3.

This table presents the matrix analysis of epistemological differences across four interaction patterns based on the three-dimensional epistemological framework.

4.2.1.

Epistemological Differences in LLM Interactions

A comparative analysis of the four interaction patterns, presented in Table

3

using a matrix method

(Groenland,

2018

)

, reveals notable epistemological differences.
This analysis, informed by the previously established three-dimensional epistemological framework

(Hofer and Pintrich,

1997

; Aksan,

2009

; Mokhtari,

2014

; Huang

et al.

,

2023

)

described in §

3.3.2

, considers the perceived

nature

of knowledge (absolute vs. evolving), the primary

source

of knowledge (external authorities vs. active construction), and the criteria for

justifying

knowledge.

The contrast between proactive and receptive interaction patterns revealed fundamentally different mental models of sensemaking.
Proactive patterns, marked by actively seeking diverse perspectives and critical reflection, facilitate dynamic knowledge construction and validation through iterative interaction and exploration.
While this iterative exploration promotes deep understanding, it creates a practical challenge of efficiently verifying the extensive information generated through LLM interactions (see

C1

in §

4.2.2

).
In contrast, receptive patterns rely heavily on external authority and a more passive way of engaging, which limits effective knowledge construction with LLMs. As a result, fostering meaningful interactivity and learner agency (see

C2

in §

4.2.2

) becomes particularly challenging when users treat LLMs primarily as information sources rather than interactive learning partners.

Participants exhibiting validation-seeking and content-focused interaction patterns occupy an interesting middle ground.

While they value accuracy and structure, their strategic use of LLMs for targeted information suggests an openness to expanding their understanding beyond curated materials.

However, these patterns tend to emphasize alignment with authoritative sources and validation of prior knowledge. They may inadvertently limit learners’ exploration of diverse perspectives and create a gap between their desire for comprehension and reliance on established sources (see

C3

in §

4.2.2

).

4.2.2.

Challenges in LLM Interactions

Our analysis of epistemological differences across different LLM interaction patterns, integrating survey (see Fig.

5

) and semi-structured interview results, reveals unique challenges participants encountered during LLM-mediated self-learning.

C1

.

Inefficient Information Verification.

Across various

interaction patterns

, effectively and efficiently verifying the accuracy of LLM-generated information poses a significant challenge

(Cucuiat and Waite,

2024

)

.

Participants with proactive and validation-seeking interaction patterns find the validation process time-consuming and cumbersome, hindering their ability to assess the reliability of AI’s responses.
For instance, P9 commented, “

ChatGPT sometimes fabricates information, so I have to verify its explanations against multiple sources, which can takes longer than finding the answer myself.

”

C2

.

Limited Interactivity and Agency.

The interactive nature of LLMs presents challenges to self-learners accustomed to more structured or passive approaches

(Zamfirescu-Pereira

et al.

,

2023

)

.

Participants showing content-focused interaction patterns,

find that the lack of seamless integration between LLM responses and their preferred learning materials diminishes their sense of control over the learning process.

As P11 noted, “

Switching between the chat and readings is frustrating. Especially when the AI’s responses do not align with my learning goals–it just throws me off track.

”

Similarly,

participants with receptive tendencies

struggle to adapt to the open-ended nature of interacting with LLMs, finding it difficult to formulate effective queries and manage the volume of information received.

As P24 explained, “

Asking good questions is challenging in itself–you need to know the topic already. Helpful prompts from the AI would be great; otherwise, I am lost.

”

C3

.

Confirmation Bias.

Participants showing validation-seeking preferences are particularly vulnerable to confirmation bias.

They use LLMs to reinforce pre-existing beliefs rather than critically examining alternative perspectives or challenging their own assumptions.

For example, P22 admitted, “

I sometimes keep tweaking my prompts until ChatGPT gives me what I expected, even if that means ignoring other perspectives.

”
This behavior underscores how unstructured LLM interactions can narrow learning pathways. It highlights the need for systematic query approaches and metacognitive support to help learners examine their thinking processes and develop more comprehensive exploration strategies.

Figure 5.

A substantial portion of participants reported challenges using the LLM-supported chatbot for self-study, particularly with question formulation, response clarity, and the achievement of deep understanding.

4.2.3.

Design Requirements

The challenges uncovered in our formative study directly informed the five key design requirements for our LLM-supported self-learning system:

DR1

.

Validated Information and Validation Support (

C1

).

Systems must provide readily validated information or offer user-friendly methods for verifying the accuracy of LLM outputs.
This includes facilitating precise cross-referencing between LLM responses and the original source material.

DR2

.

Smooth Onboarding and Reduced Cognitive Load (

C2

).

Learners expressed a preference for LLMs to autonomously generate systematic follow-up questions, minimizing the initial interaction effort. This approach reduces the barrier to entry and enhances self-study efficiency by reducing the time spent crafting prompts.

DR3

.

Sustained Engagement and Deepened Exploration (

C2

).

Engaging and fluid dialogue is crucial to maintaining users’ engagement while guiding them toward a deeper understanding of new concepts. Supporting flexible, multi-round dialogues enables learners to progressively deepen their understanding of concepts through both horizontal and vertical exploration.

DR4

.

Fostering Openness and Multidimensional Thinking (

C3

).

Comprehensive knowledge acquisition requires learners to consider multiple perspectives and generate diverse hypotheses.
A well-designed taxonomy of queries, informed by a robust epistemological framework, can effectively support this type of multifaceted thinking.

DR5

.

Facilitating Metacognition and Reflection (

C3

).

Systems should support metacognition by structuring the interaction process and encouraging reflection on learning progress and comprehension levels.
Recording and visualizing the user’s chain-of-query can enhance knowledge retrieval efficiency and help identify potential gaps in understanding.

5.

CausaDisco

We present

CausaDisco

, an epistemologically-informed system designed to enhance sensemaking during self-learning.
Carefully designed features enable

CausaDisco

to encourage user interaction with an LLM chatbot to enhance sensemaking, fulfilling the design requirements outlined in §

4.2

.
In the following sections, we illustrate

CausaDisco

’s interface design, highlighting its integration of the epistemological framework and prompt engineering techniques.

5.1.

Interface Design

CausaDisco

integrates four coordinated views to streamline the self-learning process within complex educational content.
The

Embedded Content View

(Fig.

6

A) provides users with access to the original learning materials, allowing immediate verification and reference (

DR1

).
Complementing this, the

Concept Graph View

(Fig.

6

B) visually maps out the core concepts and their interrelationships, clearly and concisely representing the structure of the material (

DR2

).
The interactive

Q&A Conversation View

(Fig.

6

C) enables users to interact with the chatbot and to choose from a range of suggested follow-up questions, fostering a deeper exploration of specific topics (

DR3, DR4

).
Lastly, the

Tree Map View

(Fig.

6

D) organizes the user’s inquiry process into a logical framework, facilitating easy retrospection and effectively tracking their learning progression (

DR5

).

Figure 6.

System Overview: The interface features four main views: A)

Embedded Content View

; B)

Concept Graph View

; C)

Q&A Conversation View

; D)

Tree Map View

.
The backend comprises two main components: (E1) The Four Causes principles, which inform the prompt engineering for generating follow-up questions (FQ); and (E2) A logging function that records the user’s chain of queries (E3) during interactions with the LLM chatbot.

5.1.1.

Embedded Content View

The

Embedded Content View

(Fig.

6

A) has two sections.
The right-hand side displays the original learning materials, including hierarchical headings, author information, publication date, main text, and illustrations (when available) (Fig.

6

A2).
The left-hand side provides an interactive navigation bar for users to locate sections of interest swiftly and can be hidden to maximize screen space (Fig.

6

A1).
This design facilitates users in verifying the authenticity or accuracy of LLM responses (

DR1

).
They can effortlessly use the navigation bar to pinpoint the pertinent section in the original text to cross-reference information and delve into more details.

Justifications:

The design of the

Embedded Content View

prioritizes text-based learning resources, as informed by our formative study findings.
This focus on text enables seamless AI integration without disrupting learning flow while allowing efficient content parsing for accurate, context-aware AI responses. Text-based content also offers flexible presentation and interaction, accommodating diverse

interaction patterns

.

5.1.2.

Concept Graph View

The

Concept Graph View

(Fig.

6

B) comprehensively summarizes the core concepts and their interrelationships within the learning materials.
We employ a semi-automated approach to enhance the accuracy and granularity of the concept graph.
By first identifying high-frequency keywords in the materials, we established the scope of the concept graph by incorporating a glossary of concepts provided by ChatGPT-4.
After analyzing the original texts with assistance from two domain experts in the respective fields (in our case, NFTs and semiotics), we identified four types of relationships between the key concepts: foundational prerequisites, defining traits, illustrative examples, and the influence of one concept’s functions or characteristics on another (Fig.

6

B2).
Accordingly, we created the corresponding concept graphs for different learning materials (Fig.

6

B1).
This view offers users a concise overview of the knowledge landscape, using keywords as search indices to create a straightforward starting point for interacting with the LLM (

DR2

).

5.1.3.

Q&A Conversation View

The

Q&A Conversation View

(Fig.

6

C) equips users with epistemologically-informed follow-up questions during their interactions with the LLM chatbot.
This view autonomously generates four alternative follow-up questions immediately after the chatbot responds to the previous query (Fig.

6

C2).
These follow-up questions support the deep exploration and comprehension of the concepts discussed, enriching the experience beyond the typical single-turn conversation (

DR3

).
Users can interact with any of the automatically generated follow-up questions by clicking on them to obtain a response from the LLM chatbot.
They can also customize a question by clicking the “modify” button next to it.
This button allows them to edit the question in the input box before submitting it.

The follow-up questions are crafted to guide users in exploring the topic from multiple perspectives (

DR4

), fostering a more holistic understanding. The underlying mechanism for generating these questions is discussed in detail in §

5.2

.
This view also logs all dialogue interactions between users and the LLM chatbot (Fig.

6

E2).

5.1.4.

Tree Map View

The

Tree Map View

(Fig.

6

D) dynamically visualizes users’ query histories, allowing for a quick review of their learning progress and facilitating timely adjustments to their learning strategies (

DR5

).
This view is characterized by its cross-view interactions. When users interact in the

Q&A Conversational View

, the

Tree Map View

simultaneously and automatically records the “parent” questions they ask.
In a hierarchical node-link format, it presents the “child” follow-up questions selected by users (Fig.

6

E3).
The query tree guides users to progressively explore the topic in depth until a new question is posed, indicating the end of the learning session for the previous parent topic.
At this point, the view automatically creates a new tree map for the new parent topic.

Users can freely expand or collapse branches associated with the same parent topic to adjust the level of detail shown in the tree map.
They can also adjust the tree map’s size and position through scrolling and panning.
This view preserves all generated tree maps, assisting users in managing their interaction records and reviewing their logical structures, supporting reflection on the learning process, and adjusting their question strategies.

Justifications:

We also considered using a table format to display users’ query histories.
However, feedback from participants during our formative study suggested that graphs are more intuitive than tables, as they better illustrate the hierarchical relationships between questions.

5.2.

Prompt Engineering

This section details the process of generating the follow-up questions outlined in §

5.1.3

. We discuss the epistemological framework we selected, the rationale behind its choice, and how it is incorporated into our prototype.

5.2.1.

Leveraging the Epistemological Framework

As discussed in §

4.2

, we aim to design follow-up questions that facilitate both comprehensive exploration and multidimensional thinking. Assigning the LLM, ChatGPT in our case, the role of a tutor seems appropriate; however, this poses the challenge that the generated questions might merely be relevant without addressing the user’s understanding gaps. Furthermore, ChatGPT’s tendency to mimic expert vocabulary

(Zollman

et al.

,

2023

)

may be unsuitable for our application scenario. Therefore, we equip ChatGPT with a

structured framework

to guide its output, ensuring the follow-up questions provide effective cognitive support for self-learners.

Given that the follow-up questions are designed to facilitate a comprehensive exploration of the topic, incorporating this

framework

into the prompt should yield questions that span all categories we summarize in Table

2

.
Furthermore, our approach needs to be domain-agnostic and easily integrated into prompts.
Consequently, we opted for an established theory that ChatGPT understands and incorporated it as a guiding term in our prompts.

After exploring various options, we found that popular learning theories from behaviorism or cognitivism were inadequate for our needs.
Additionally, studies on student mental models often focus on specific domains

(Fratiwi

et al.

,

2020

; Taylor

et al.

,

2003

)

.
To address this, we turned to epistemology, which centers on acquiring knowledge and understanding reality. Epistemological theories have long been integrated into pedagogical practices to foster critical thinking and active engagement in exploring the unknown

(Duschl

et al.

,

1990

; Macallister,

2012

; Kotzee,

2018

)

.
Among various epistemological theories, Aristotle’s

Four Causes

framework (

Material

,

Formal

,

Efficient

, and

Final

Causes)

(Hocutt,

1974

)

stands out for its ability to refine cognitive processes

(Pérez and Ziemke,

2007

)

and improve educational outcomes

(Taylor and Sondermeyer,

2023

)

.

Specifically, this framework offers a comprehensive analytical lens for understanding the essence of any entity, covering all question types summarized in Table

2

.

We used a deductive method to map the connections between our query taxonomy and the

Four Causes

framework (Table

4

). The material cause naturally corresponds with the “components of concepts” category, which refers to the fundamental elements constituting a concept. The formal cause, which describes the distinct form or arrangement defining an object, aligns with the “attributes of concepts” category. Additionally, the formal cause includes “co-existent concepts,” helping clarify the characteristics that distinguish a concept from others. The efficient cause, representing the driving force behind an entity’s existence or change, aligns with both “realistic application” and “development of concepts,” as these categories explore the motivations and mechanisms of change. Lastly, the final cause, which denotes the purpose or intended outcome of an entity, maps to the “significance of concepts” and “real-world consequences” categories

(Falcon,

2006

; Cohen and Reeve,

2000

; Charlton and others,

1983

)

.

Table 4.

The mapping between the epistemological framework, i.e.,

Four Causes

, and the query taxonomy identified from the dialogue data generated by participants during the self-learning task.

Four Causes

Original Definition

Query Types

Examples

Material Cause

“

That out of which a thing comes

to be and which persists.

”

Components

of Concepts

“Why must NFTs be involved with blockchain?”

“Can users mint NFTs without any assets?”

Formal Cause

“

The form or the archetype,

i.e. the statement of the essence.

”

Attributes

of Concepts

“What is the meaning of non-fungible?”

“Do NFTs protect copyrights?”

Co-existent

Concepts

“What are the benefits of NFTs compared to cryptocurrencies?”

“What are the relationships between NFTs, ETH, and blockchains?”

Efficient Cause

“

The primary source of change and

the various factors that contribute

to that change.

”

Realistic

Application

“How are Non-fungible tokens (NFTs) regulated?”

“What are the uses of NFTs beyond investigation and collection?”

Development

of Concepts

“How was the first NFT created?”

“What is the outlook for NFTs?”

Final Cause

“

In the sense of end or ‘that for

the sake of which’ a thing is done.

”

Significance

of Concepts

“What is the importance of NFTs?”

“Why do we need NFT as a newly-emerging technology?”

Real-world

Consequences

“What are the benefits of owning an NFT?”

“What is the significance of NFTs in the current capitalist society?”

5.2.2.

Integrating Epistemological Framework

Here we elucidate the process of generating follow-up questions using the

Four Causes

framework within our system, leveraging the capabilities of ChatGPT. It is worth noting that while a foundational understanding of the Four Causes is not a prerequisite for users, the questions are strategically designed to guide their inquiry towards more complex and contextually rich content.

In our initial evaluation, we assessed ChatGPT’s ability to comprehend and apply the Aristotelian Four Causes.
The results were promising, demonstrating ChatGPT’s general adeptness in providing illustrations for each cause in relation to specific concepts or entities. Subsequent to this analysis, we investigated the system’s proficiency in crafting follow-up questions anchored in the Four Causes framework.
ChatGPT exhibited remarkable efficiency in generating inquiries that were not only relevant to the ongoing conversation but also adeptly integrated aspects from the most recent query, thus maintaining a coherent and contextually relevant learning trajectory.

However, it was observed that ChatGPT sometimes struggled to accurately assign the appropriate cause category to each question and to prioritize certain causes over others based on the context or scenario at hand. This suggests a potential area for refinement in the underlying language model, which will be explored more comprehensively in the Discussion section (§

8.3.3

). Despite these challenges, the utilization of the generic model facilitated a structured and in-depth exploration of topics, thereby enhancing the self-learning experience by guiding learners through a more nuanced understanding of the subject matter.

To further refine the guidance provided and enhance its pedagogical value, we incorporated the

persona pattern

, as outlined by White

et al.

(White

et al.

,

2023

)

.
This methodology involves configuring ChatGPT to function as an educational tutor, adapting its way of interacting to more effectively facilitate a learning environment. Within this adapted role, ChatGPT’s responsibilities extend beyond merely answering questions; it is also engineered to proactively generate follow-up questions. For each interaction, the directive “

Provide the top four related follow-up questions based on the previous question using the four causes idea

” is employed. This strategy ensures the elicitation of a set of questions that are not only relevant but also imbued with educational significance, thus fostering a more engaging and productive self-learning experience.

5.3.

Usage Scenario

This section illustrates how Emily, a Ph.D. candidate with an interdisciplinary background, efficiently self-learned a new knowledge domain (NFTs) with

CausaDisco

.
In her previous self-study endeavors, Emily had often relied on LLM tools to help her gather and consolidate information to build a foundational understanding of new concepts.
For example, she mentioned using AskYourPDF

6

6

6

https://askyourpdf.com/zh

and Poe

7

7

7

https://poe.com/

to quickly extract information from references.
Nevertheless, crafting the right questions can be time-consuming, and although some LLM tools offer follow-up questions, they often lack relevance or do not align with her interests.
As a result, these tools typically only provide hints or keywords, leaving her to rely on traditional search engines or to read original materials for in-depth learning.

Upon adopting

CausaDisco

, Emily’s learning process evolved. She first examined the

Concept Graph View

in the upper right corner to gain an overview of the core concepts.
Then, with the help of the navigation bar, she quickly reviewed the

Embedded Content View

in the upper left corner.
As a beginner in NFTs, Emily zeroed in on key concepts and sought clarification through the

Q&A Conversation View

, asking, “

Can you explain what fungible means in one sentence?

”
After reading the reply and realizing that each NFT’s price varies–a topic she found intriguing–she inquired further, “

How do NFT creators determine their products’ prices?

”
After receiving this response, she engaged in five turns of interaction with the AI chatbot on the topic.
Under the guidance of the follow-up questions, she conducted an in-depth exploration of NFT pricing and the influencing factors, achieving a comprehensive understanding.

By reviewing the

Tree Map View

, Emily noted that she had explored the factors affecting NFT pricing from various angles, including the personal expectations of artists, the volatility of cryptocurrency prices, and the pricing differences between international and local markets.
She also understood how to address price volatility.
Satisfied with the depth and breadth of her exploration, Emily then started on a new topic—how NFTs influence society—and pursued deep self-learning through a workflow similar to the one she had previously followed.

6.

Evaluation

We conducted a within-subject study to evaluate whether

CausaDisco

enhances knowledge sensemaking during self-learning.
Specifically, we aimed to investigate the following three research questions (RQs):

RQ1:

How does

CausaDisco

foster interactivity with the LLM chatbot?

RQ2:

How does

CausaDisco

facilitate sensemaking in self-directed learning?

RQ3:

How do users perceive the usefulness and experience of interacting with

CausaDisco

?

6.1.

Participants

Our study involved 36 participants (19 male, 17 female) ranging in age from 20 to 36 years old (M = 26.31, SD = 3.01).
We recruited participants through convenience and snowball sampling, using criteria similar to those of our formative study.
Participants had diverse educational backgrounds, with 13 holding bachelor’s degrees, 15 holding master’s degrees, and 8 holding doctorates.
Most participants (n = 25) specialized in STEM fields, while others came from the humanities and social sciences (n = 5), interdisciplinary studies (n = 3) and medical fields (n = 2).
Participants also had varying levels of academic training experience (more than five years: n=11; three to five years: n=13; fewer than three years: n=12). Familiarity with LLM chatbots ranged from extremely familiar (n = 3) and very familiar (n = 15) to somewhat familiar (n = 11), somewhat unfamiliar (n = 6) and neutral (n = 1).

6.2.

Procedure

We employed a within-subject experimental design in which participants independently completed two self-learning tasks, each focused on a distinct concept: NFTs

(Schulze,

2024

)

and semiotics

(StudySmarter,

2024

)

.
To minimize order bias, we counterbalanced the order of exposure to our system,

CausaDisco

, and a baseline condition (using participants’ preferred tools), resulting in 4 (= 2 x 2) conditions

(Guo

et al.

,

2023

; Suh

et al.

,

2023

)

.
Participants were encouraged to engage in active learning through conversational interaction with the LLM chatbot in both conditions (see Table

7

in Appendix

B

).

The study began with participants providing informed consent and completing a pre-study survey with demographic questions. They then engaged in two self-learning tasks, using either

CausaDisco

or their chosen baseline tools to explore one of the assigned topics. The procedure for self-learning tasks involved four distinct sessions:

(1)

Introduction:

A brief overview of the research background and experimental protocols.

(2)

Pre-Task Exercise:

Participants were given 10 minutes for free exploration of the assigned topic.

(3)

Self-Learning Tasks:

Participants engaged in each self-learning task for 20 minutes using either

CausaDisco

or their chosen baseline tools.

(4)

Assessment:

Participants engaged in a 15-minute quiz comprising three multiple-choice and three open-ended questions. This quiz, developed in collaboration with domain experts in NFTs and semiotics, assessed the participants’ understanding of the presented concepts.
Following the quiz, participants completed a post-task survey using a seven-point Likert scale to evaluate how each condition supported their sensemaking during self-learning.
The session concluded with semi-structured interviews where participants compared

CausaDisco

and the baseline condition, providing feedback on the system design.

The entire experiment lasted approximately 90 minutes.
The 20-minute duration for each self-learning task was determined based on observations from a pilot study with team members.

Figure 7.

This figure presents participants’ subjective ratings (seven-point Likert scale) comparing

CausaDisco

and a baseline on interactivity and sensemaking.

6.3.

Measures

We evaluated participants’ perceptions of interactivity, sensemaking, and the usability and design of

CausaDisco

. All measures were rated using a seven-point Likert scale (1 = Strongly Disagree, 7 = Strongly Agree).

6.3.1.

Interactivity Measures

To assess how

CausaDisco

enhanced interactivity, we analyzed both objective measures and subjective participant perceptions.
We began by analyzing the dialogue records generated during the self-learning tasks, counting and comparing the number of dialogue turns between participants and LLMs in both the

CausaDisco

and baseline conditions.
This analysis provides an objective measure of interaction frequency.

To evaluate participants’ subjective experiences of interactivity, we examined two key dimensions: perceived

engagement

and

efficiency

(Fig.

7

).

Each dimension comprised two subdimensions. Engagement was assessed by participants’ perceived logical coherence of system-generated follow-up questions (

Q1

) and their perceived inspiration for exploring new knowledge domains (

Q2

).
Efficiency was evaluated by participants’ perceived ability to efficiently accomplish self-learning objectives (

Q3

) and the perceived streamlining of the learning process (

Q4

).

6.3.2.

Sensemaking Measures

Building upon the sensemaking model proposed by Pirolli

et al.

(Pirolli and Card,

2005

)

, we assessed how

CausaDisco

facilitates sensemaking through both subjective and objective measures. For subjective evaluation, we examined two key dimensions: perceived

comprehensiveness

and

sophistication

(Fig.

7

).

Comprehensiveness was evaluated by participants’ perceived ability to obtain sufficient information with the system (

Q5

), their perceived support for comprehensive exploration of new knowledge domains (

Q6

), and their perceived acquisition of multi-dimensional perspectives during self-learning (

Q7

). Sophistication was assessed by participants’ perceived deepening of understanding within a new knowledge domain (

Q8

).

For objective evaluation, we recruited domain experts—two Ph.D. students in Fintech and two with master’s degrees in communications—to blindly evaluate participants’ quiz responses. Each participant’s final score was calculated as the average of two expert ratings. The multiple-choice questions assessed the comprehensiveness of understanding, while short-answer questions evaluated the sophistication of knowledge application.

6.3.3.

Usability and Design Measures

We evaluated how users perceived

CausaDisco

’s usability and design using the technology acceptance model

(Venkatesh and Bala,

2008

)

.
Participants rated their perceptions of the system’s ease of use, ease of learning, and intention to recommend

CausaDisco

on a seven-point Likert scale.
Furthermore, we assessed participants’ perceived intuitiveness of the system, encompassing both interface and interaction design elements.

Table 5.

Within-subject comparisons of participants’ subjective ratings in the two conditions. Measures include perceived engagement, efficiency, comprehensiveness, and sophistication. Statistical significance: *

p

<

0.05

p<0.05

, **

p

<

0.01

p<0.01

, ***

p

<

0.001

p<0.001

,

*

p

<

0.1

p<0.1

(marginally significant).

Measures

Attributes

Baseline

CausaDisco

Statistics

Mean/SD

Mean/SD

t

df

p

Sig

Engagement

Logically Coherent

5.28/1.28

6.06/0.95

3.08

35

0.001

**

Inspiring

5.31/1.31

6.08/0.65

3.08

35

0.001

**

Efficiency

Efficiency Enhancement

5.31/1.04

5.64/1.15

1.83

35

0.038

*

Streamlined Process

5.25/1.30

5.53/1.18

1.06

35

0.149

Comprehensiveness

Sufficient Information

5.69/1.12

5.69/1.31

0

35

0.500

Comprehensive Exploration

5.31/1.21

5.72/1.32

1.46

35

0.076

*

Multi-Dimensional

4.89/1.37

5.67/1.39

2.38

35

0.011

*

Sophistication

In-Depth Understanding

5.14/1.38

5.78/1.05

2.44

35

0.009

**

7.

Results

This section presents findings from our analysis of participant survey responses, prototype interaction logs, and interview feedback.
To compare quantitative data from surveys and system logs, we used paired samples t-tests between the two conditions, except when Shapiro-Wilk tests indicated deviations from normality.
In those cases, we used the non-parametric Wilcoxon Signed Rank test.

7.1.

Interactivity Enhancement (RQ1)

To evaluate

CausaDisco

’s interactivity, we assessed both perceived engagement and efficiency using a combination of objective and subjective measures (see Table

5

and Fig.

8

).

7.1.1.

Engagement

Participants’ ratings of logical coherence are significantly higher in the

CausaDisco

condition (

M

=

6.06

M=6.06

,

S

​

D

=

0.95

SD=0.95

) than in the baseline condition (

M

=

5.28

M=5.28

,

S

​

D

=

1.28

SD=1.28

,

p

<

.01

p<.01

).
Similarly, participants’ ratings indicated significantly higher inspiring of the

CausaDisco

condition (

M

=

6.08

M=6.08

,

S

​

D

=

0.65

SD=0.65

) than the baseline condition (

M

=

5.31

M=5.31

,

S

​

D

=

1.31

SD=1.31

,

p

<

.01

p<.01

).
Aligned with the self-rating results, participants engaged in significantly more dialogue turns in the

CausaDisco

(

M

=

6.97

M=6.97

,

S

​

D

=

2.22

SD=2.22

) than in the baseline condition (

M

=

4.08

M=4.08

,

S

​

D

=

1.86

SD=1.86

,

p

<

.01

p<.01

).

Qualitative data from participant interviews corroborates these quantitative findings.
Many participants were impressed by

CausaDisco

’s epistemologically informed follow-up questions, describing them as “

more logically structured, providing a clear train of thought

” (P35) and “

more human-like with high-quality

” (P34).
Furthermore, participants reported that

CausaDisco

facilitated the discovery of new points of interest.
For instance, P25 stated, “

The follow-up questions and concept graph of CausaDisco are eye-openers! It is like having a curiosity engine–the more I use it, the more I want to keep chatting with the AI.

”
The combination of quantitative and qualitative results provides compelling evidence for

CausaDisco

’s effectiveness in enhancing user engagement with LLM chatbots.

7.1.2.

Efficiency

Participants reported significantly higher efficiency enhancement in the system condition (

M

=

5.64

,

S

​

D

=

1.15

M=5.64,SD=1.15

) compared to the baseline condition (

M

=

5.31

,

S

​

D

=

1.04

,

p

<

.05

M=5.31,SD=1.04,p<.05

). Participants in the system condition, on average, rated the streamlined process 5.53 (

S

​

D

=

1.18

SD=1.18

), higher than the baseline condition (

M

=

5.25

,

S

​

D

=

1.30

M=5.25,SD=1.30

), but there are no significant differences.

The results underscore the potential of our system to enhance user efficiency in a significant yet implicit manner.
Although participants reported perceiving an increase in their efficiency, this improvement was not attributed to a more streamlined process.
As P4 remarked, “

The follow-up questions of CausaDisco are spot-on and boost my self-learning efficiency. I learn faster without having to rack my brain thinking up questions myself. It is a real time-saver.

”
This suggests that our system may boost participants’ efficiency by delivering precisely the questions they require at the moment.

However, the underlying logic and mechanism responsible for generating the follow-up questions are not fully comprehensible to participants.
For instance, P29 noted, “

I often pause to consider the underlying logic of the follow-up questions and the learning trajectory in the tree map. The deeper I go with the questions, the more I find myself taking these little moments to reflect.

”
This uncertainty may lead participants to doubt the effectiveness of the questions in streamlining their self-learning.
This hypothesis is corroborated by the lack of significant findings in streamlined process, for future research and system improvements to better support users’ self-learning.

Figure 8.

Within-subject comparison of participants’ subjective ratings of key measures (engagement, efficiency, comprehensiveness, and sophistication) across two conditions.

Error bars represent between-subjects standard error

; asterisks indicate statistical significance from paired t-tests (*p ¡ 0.05, **p ¡ 0.01). The difference in perceived comprehensive exploration, though not statistically significant, approached significance and is highlighted in gray.

7.2.

Sensemaking Support (RQ2)

To assess how

CausaDisco

facilitates sensemaking, we compared participants’ perceived

comprehensiveness

and

sophistication

when processing information under the two conditions (see Table

5

and Fig.

8

).

7.2.1.

Comprehensiveness

We found no significant differences between conditions regarding sufficient information (the system condition:

M

=

5.69

,

S

​

D

=

1.31

M=5.69,SD=1.31

; the baseline condition:

M

=

5.69

,

S

​

D

=

1.12

;

p

>

.05

M=5.69,SD=1.12;p>.05

). However, participants reported marginally significantly higher comprehensive exploration in the system condition (

M

=

5.72

,

S

​

D

=

1.32

M=5.72,SD=1.32

) compared to the baseline condition (

M

=

5.31

,

S

​

D

=

1.21

,

.05

<

p

<

.10

M=5.31,SD=1.21,.05<p<.10

). They also self-reported higher multi-dimensional ratings in the system condition (

M

=

5.67

,

S

​

D

=

1.39

M=5.67,SD=1.39

) compared to the baseline condition (

M

=

4.89

,

S

​

D

=

1.37

,

p

<

.05

M=4.89,SD=1.37,p<.05

). Participants’ objective quiz scores in the system condition are, on average, 2.00 (

S

​

D

=

0.93

SD=0.93

), a little lower than in the baseline condition (

M

=

2.25

,

S

​

D

=

0.89

M=2.25,SD=0.89

), but there is no significant difference (

p

>

.05

p>.05

).

The results demonstrate that our system did successfully deliver multi-dimensional information to users. However, the scope of the information was constrained by the specific study materials provided. Consequently, participants expressed mixed opinions on whether the new information was sufficient for a fully comprehensive understanding.

In line with our goal of offering diverse perspectives on unfamiliar topics,

P30 shared, “

Compared to my usual tools, CausaDisco’s follow-up questions focus on key ideas while encouraging broader thinking. Along with the concept map, which gives me a big-picture view of concept connections, CausaDisco helps me grasp new information in a more well-rounded way.

”
This achievement also aligns with our primary objective of providing a comprehensive self-learning environment.

However, we encountered limitations in conclusively proving that our system offers sufficient information.
For instance, P9 commented, “

It seems all info CausaDisco gives, including the follow-up questions and responses, comes straight from the study materials. That’s good for accuracy, but adding some outside facts with sources would help my learning even more.

”
This perceived lack of information likely arises from the contrasting exploration approaches between the two conditions.

CausaDisco

guides users through a specific article, generating precise, content-based follow-up questions and responses for accuracy.
Conversely, the baseline condition allows unrestricted exploration.
This key difference–

CausaDisco

’s focused, article-centric approach versus the baseline’s open-ended exploration–likely created an impression of limited information within our system.

7.2.2.

Sophistication

Participants reported significantly higher in-depth understanding ratings in the system condition (

M

=

5.78

M=5.78

,

S

​

D

=

1.05

SD=1.05

) compared to the baseline condition (

M

=

5.14

M=5.14

,

S

​

D

=

1.38

SD=1.38

,

p

<

.01

p<.01

).
Subjective quiz scores were also higher in the system condition (

M

=

5.06

M=5.06

,

S

​

D

=

0.98

SD=0.98

) than in the baseline condition (

M

=

4.56

M=4.56

,

S

​

D

=

1.12

SD=1.12

,

.05

<

p

<

.1

.05<p<.1

), although this difference was only marginally significant.

The significant improvement in perceived in-depth understanding, coupled with the trend towards better quiz performance, indicates that

CausaDisco

is likely promoting a more sophisticated exploration during self-learning.
As P1 put it, “

CausaDisco, especially the follow-up questions and the tree map, lets me really dive deep into a topic.

”
This aligns with our goal of enabling users to seamlessly connect and integrate knowledge through dialogue-driven interaction.

It is important, however, to acknowledge that subjective perceptions of understanding do not always align with objective outcomes. Future research should further evaluate the system’s efficacy in not only deepening conceptual understanding but also supporting the practical application of acquired knowledge.

Figure 9.

This figure illustrates participants’ subjective ratings of usability and design for

CausaDisco

, which were measured using a seven-point Likert scale.

7.3.

Usability and Design (RQ3)

We also employed a seven-point Likert scale and conducted semi-structured interviews to assess the usability, interface, and interaction design of

CausaDisco

(see Fig.

9

).

7.3.1.

Usability

More than half of the participants strongly agreed that our system is easy to learn and use (

Q9-10

).
During the five- to ten-minute introduction, nearly all could understand the main functions of each view of the system and quickly became familiar with the interactive operations through free exploration.
Furthermore, most participants were willing to recommend

CausaDisco

to other self-learners (

Q11

).
Only two participants (P23 and P8) were less inclined to do so, mainly because they desired more flexible and personalized interaction.
S4 explained: “

I hope the system can further enhance the design and functionality of the tree map by allowing users more freedom to annotate or organize, which could be more helpful

.”

7.3.2.

Design

Participants generally found the interface design to be intuitive (

Q12

).
The

Q&A Conversation View

and the

Concept Graph View

were particularly well-received.
For instance, P1 and P4 mentioned, “

The Concept Graph helps provide me with a big picture and clear direction for active self-learning

.”
Additionally, several participants pointed out that the

Tree Map View

helped them trace their learning journey and adjust their query strategies in a timely manner.
P7 said, “

The Tree Map View is highly beneficial as it facilitates the efficient organization of the questions I pose, aiding in cognitive training to improve question formulation

.”
Furthermore, some participants noted that the

Embedded Content View

alleviated their distrust of the AI chatbot, enabling them to validate original learning materials efficiently.
As P5 mentioned, “

The Embedded Content View with navigation bar also assists me in quickly locating relevant sections

.”

In line with the positive feedback received on the interface design, most participants deemed the interaction design of the system (

Q13

) as both “

streamlined and essential

.”
Despite this general consensus, two participants, P22 and P24, offered slight critiques.
P22, in detail, stated, “

While the system is generally user-friendly, a thorough introduction is still indispensable.
Lacking such guidance, one might spend considerable time figuring out the system independently

.”

8.

Discussion

This research aims to integrate epistemological frameworks with LLMs to enhance self-learning. We introduce

CausaDisco

, an interactive system designed to foster comprehensive exploration and holistic understanding through reflective multi-turn dialogue. Our study investigates

epistemological differences across different interaction patterns

. This leads to the development of a robust framework for generating follow-up questions to enhance learners’ sensemaking processes.
Our work contributes to the growing body of research on LLM-based educational agents and human-LLM interaction within the HCI community

(Suh

et al.

,

2023

; Zheng

et al.

,

2024b

; Hu

et al.

,

2024

; Gao

et al.

,

2024

)

, aligning with advocates for inclusive, efficient LLM applications in education

(Shah and Bender,

2022

; Fiannaca

et al.

,

2023

; Zamfirescu-Pereira

et al.

,

2023

)

.
In the following section, we discuss the design implications derived from developing and evaluating

CausaDisco

, its limitations, and potential directions for future work.

8.1.

Inclusive Design for Diverse LLM Interaction Patterns

The advent of LLMs has catalyzed a paradigm shift in education, promising customized learning experiences tailored to individual needs and preferences

(Mogavi

et al.

,

2024

)

.
Achieving this, however,

requires “

dropping an LLM-based agent as a one-size-fits-all solution

”

(Shah and Bender,

2022

)

and understanding how the

diverse interaction patterns between learners and LLMs can inform the design of inclusive and adaptive AI tools for self-learning.

Our qualitative study,

informed by the ICAP framework for classifying learner engagement

(Chi,

2009

; Chi and Wylie,

2014

)

,

revealed four distinct LLM interaction patterns:

proactive

,

validation-seeking

,

content-focused

, and

receptive

.

These patterns represent different engagement strategies and underlying epistemological schemas that learners may adopt, varying in interaction frequency, knowledge-seeking breadth, and strategic diversity. The choice of strategy appears to be influenced by contextual factors, prior learning experiences, and epistemological beliefs about the nature, source, and justification of knowledge.
We observed that when learners adopted proactive strategies, they demonstrated higher comfort levels with AI-assisted exploration. Conversely, when employing receptive or content-focused strategies–particularly common among those familiar with traditional, input-oriented pedagogy–their LLM engagement was less exploratory. In our LLM-mediated learning context, more passive engagement strategies correlated with lower overall interactivity and diminished quiz performance.
These findings highlight the need for more inclusive LLM-based learning tools that effectively support learners across diverse engagement strategies.

Designing educational tools that consider the spectrum of

interaction patterns

is essential.

Learners exhibiting proactive tendencies

may thrive with LLM-based tools that afford expansive application scenarios, diverse interaction modalities, and opportunities for verifying LLM responses.
Conversely, simply providing powerful affordances may prove insufficient for learners

who tend to interact with LLMs in a more restricted manner, particularly those

less familiar or comfortable with AI-driven learning environments.
Extant research suggests that these learners may benefit from carefully designed incentives, structured onboarding experiences, or “ease-in” approaches that gently encourage interaction and foster a sense of self-efficacy

(Zamfirescu-Pereira

et al.

,

2023

; Fiannaca

et al.

,

2023

)

.
Without such considerations, these learners risk being underserved, unable to fully realize the transformative potential of LLMs.

It is critical to underscore that this analysis does not aim to privilege one

interaction pattern

or epistemological orientations over another.
Rather, it highlights the imperative of acknowledging and accommodating the heterogeneity of learners’ cognitive processes and prior experiences when designing for LLM-mediated learning environments.
Effective, usable, and inclusive LLM-based tools must be intentionally designed to cater to this diversity, ensuring that all learners, regardless of their prior experience with AI or their preferred learning approaches, can engage with these tools confidently, comfortably, and meaningfully.

8.2.

Design Implications

8.2.1.

Supporting Knowledge Construction via Mental Models

The transformative potential of LLMs in education can be significantly enhanced by adopting an epistemological approach to prompt engineering. While current LLM-based educational tools show promise in assisting self-learners, their effectiveness and generalizability are often constrained by domain-specific designs and assumptions about user expertise

(Sheng

et al.

,

2023

; Chen

et al.

,

2024b

)

.
As LLM tools increasingly shape the educational landscape, it is crucial to democratize access and expand their applicability, particularly in providing cognitive support.
To achieve this, developers and designers should consider integrating well-established epistemological frameworks into LLM design.

Our findings, aligning with educational research, highlight the benefits of incorporating systematic guiding mechanisms in LLM interactions

(Peters,

2000

; Bråten,

2010

; Krasmann,

2020

)

.
By encouraging learners to articulate their reasoning, explore diverse perspectives, and integrate new information with existing knowledge, LLM-based educational tools can inspire deeper cognitive engagement.
Grounded in learning sciences and cognitive principles

(Aksan,

2009

; Mokhtari,

2014

; Baker and Anderman,

2020

)

, this approach can enhance critical thinking and facilitate robust knowledge construction.

8.2.2.

Fostering Interactivity for Effective LLM Education

Current LLM-based educational tools, primarily dialogue-driven, face challenges in maintaining user engagement and reducing dropout rates.

This challenge mirrors a persistent issue in pedagogical research: the difficulty of designing activities that foster constructive and interactive engagement

(Chi

et al.

,

2018

)

.

While the HCI community has explored promising directions through teacher-guided approaches

(Kumar

et al.

,

2023

)

and multi-round dialogues

(Hu

et al.

,

2024

)

, our findings indicate the need for more comprehensive strategies to support LLM-mediated sensemaking in self-learning.

Our research demonstrates that systematic follow-up questions, enhanced by concept graphs, effectively stimulate exploration and promote deeper understanding.
However, to cultivate a truly immersive and engaging self-learning environment, we must consider diverse interaction patterns.
This can be achieved by, for instance, embodying LLMs with distinct personas

(Liu

et al.

,

2024b

)

, incorporating interactive visualizations

(Gao

et al.

,

2024

)

, and integrating gamification elements

(Lacerda

et al.

,

2024

)

.
Such approaches not only cater to diverse

interaction patterns

but also mitigate cognitive load for AI novices.

8.2.3.

Managing Learners’ Sensemaking Process

To create more effective learning experiences, designers and developers could address the distinct stages of the sensemaking process in learning.
This process, which encompasses information retrieval, concept comprehension, knowledge integration, and application, engages both inductive and deductive reasoning

(Pirolli and Card,

2005

; Suh

et al.

,

2023

; Zheng

et al.

,

2024b

)

.

Our research demonstrates that visualizing learners’ chains-of-query and promoting metacognition enhance understanding for self-learners.
Based on these findings, we propose developing responsive and adaptive LLM-based educational tools that explicitly support and display various sensemaking stages.
For instance, implementing a meta-layer to track and visualize the learner’s granular learning steps can further boost self-awareness and learning efficacy. Such features would enable learners to monitor their progress and adjust their strategies more efficiently.

8.3.

Limitations and Future Work

Although our study comprehensively investigated participants’ LLM interactions during self-learning and identified their epistemological differences to inform the development of

CausaDisco

, it is not without limitations.

8.3.1.

Method Limitations

As with many qualitative studies, our findings are limited by a modest sample size

(Hammarberg

et al.

,

2016

)

, potential researcher biases

(Galdas,

2017

)

, and the as-yet unconfirmed generalizability of our results

(Hadi Mogavi

et al.

,

2022

)

.

All of our participants–well-educated adults with substantial academic and self-learning experience–had a clear sense of “what they should know,” even though some struggled to ask specific questions.
While our sample provided valuable initial insights, the dynamics of LLM-mediated self-learning likely vary across different educational backgrounds. Learners with less formal education potentially require different forms of support.

Future work could therefore prioritize validating and extending our taxonomy of self-learners’ interaction patterns and epistemological differences by: (1) including larger, more diverse samples; (2) exploring a wider range of subject domains; and (3) integrating quantitative methods.

These efforts will inform the development of LLM-based tools that more effectively support learners across different educational backgrounds.

8.3.2.

Expanding Personalization

While most participants appreciated the lightweight interactions offered by

CausaDisco

, they also expressed a desire for a more customized experience.
This resonates with Butcher’s call

(Butcher and Sumner,

2011

)

for cognitive personalization technologies to support meaningful analysis and coherent understanding among learners with varying levels of prior knowledge.
Our evaluation revealed that participants saw the concept graph not merely as an informative display but as a potentially powerful navigational tool for exploring learning materials.
Thus, they envisioned more flexible interaction features that enable cross-view connections between reading materials, dialogue with the LLM, and the concept graph itself. This user-driven design approach would offer an alternative, logically structured pathway through the material, rather than organizing all keywords according to the author’s logic.
Future development will focus on enhancing interactivity, particularly within the concept graph, to create richer “revisit” channels

(Pirolli and Card,

2005

)

for personalized sensemaking, ultimately improving both comprehension and retention.

8.3.3.

Four Causes Across Diverse Disciplines

In our prototype, we employed ChatGPT to generate follow-up questions without specifying a particular cause from Aristotle’s framework. This approach leverages the inherent word associations within the language model to generate contextually appropriate questions, without assessing and understanding the differential significance attributed to each of Aristotle’s

Four Causes

within these diverse fields.

Notably, not every discipline inherently involves all four causes. The significance of each cause can vary greatly across different domains, with some causes holding negligible interest or relevance in certain knowledge fields.
For instance, Aristotle, in his

Metaphysics

, posited that lunar eclipses lack a final cause and, strictly speaking, do not pertain to “matter” in the conventional sense. Similarly, within undergraduate physics education, the hypothetical final cause of gravity, despite its speculative existence, generally does not engage the curiosity or concern of students. This disinterest in the final cause is so prevalent within scientific discourse that some asserted, “

Math rejects the final cause.

”
Therefore, to achieve a more targeted and nuanced guidance, a refinement of the language model is imperative. This involves integrating a sophisticated understanding of the varying weights and significances of the

Four Causes

across different disciplines into the model, thereby enhancing the precision and relevance of the generated follow-up questions.

9.

Conclusions

This work investigates the design of inclusive LLM-based educational agents that support diverse interaction patterns. Our initial formative study (N=26) examined self-learners’ behaviors and mental models during LLM-mediated self-learning, identifying four distinct interaction patterns. Through an epistemological analysis of these interaction patterns, we identified three key challenges in learner-LLM interaction and derived five corresponding design requirements.

To address these challenges, we developed

CausaDisco

, a dialogue-driven LLM-based system designed to facilitate sensemaking and mental model construction when learners engage with complex information.

CausaDisco

integrates a query taxonomy derived from our formative study with the epistemological principles of Aristotle’s Four Causes into its prompt engineering process. This approach generates coherent and contextually relevant follow-up questions that promote deeper understanding. A subsequent within-subjects study (N=36) demonstrated that

CausaDisco

fostered more engaging interactions, enhanced perceived comprehension and learning depth, and provided a more intuitive and effective interface. This research advances our understanding of how LLMs can serve as effective educational agents for sensemaking in self-learning and offers valuable design implications for this emerging class of LLM-mediated learning tools.

Ethical Approval

The research involving human participants was performed in accordance with the principles of the Declaration of Helsinki. The study protocol was reviewed and approved by the Human and Artefacts Research Ethics Committee (HAREC) at the Hong Kong University of Science and Technology (HKUST), under the reference number HREP-2024-0043. Verbal informed consent was obtained from all participants included in the study.

References

R. Abdelghani, Y. Wang, X. Yuan, T. Wang, P. Lucas, H. Sauzéon, and P. Oudeyer (2024)

GPT-3-driven pedagogical agents to train children’s curious question-asking skills

.

International Journal of Artificial Intelligence in Education

34

(

2

),

pp. 483–518

.

Cited by:

§1

.

M. Akhtar, J. Neidhardt, and H. Werthner (2019)

The potential of chatbots: analysis of chatbot conversations

.

In

2019 IEEE 21st Conference on Business Informatics (CBI)

,

Vol.

01

,

pp. 397–404

.

External Links:

Document

Cited by:

§1

.

N. Aksan (2009)

A descriptive study: epistemological beliefs and self-regulated learning

.

Procedia-Social and Behavioral Sciences

1

(

1

),

pp. 896–901

.

Cited by:

§3.3.2

,

§3.3

,

§4.2.1

,

§8.2.1

.

M. Aliannejadi, L. Azzopardi, H. Zamani, E. Kanoulas, P. Thomas, and N. Craswell (2021)

Analysing mixed initiatives and search strategies during conversational search

.

In

Proceedings of the 30th ACM International Conference on Information & Knowledge Management

,

pp. 16–26

.

Cited by:

§3.3

.

A. R. Baker and L. H. Anderman (2020)

Are epistemic beliefs and motivation associated with belief revision among postsecondary service-learning participants?

.

Learning and Individual Differences

78

,

pp. 101843

.

Cited by:

§1

,

§2.2

,

§8.2.1

.

M. Bernabei, S. Colabianchi, A. Falegnami, and F. Costantino (2023)

Students’ use of llm in engineering education: a case study on technology acceptance, perceptions, efficacy, and detection chances

.

Computers and Education: Artificial Intelligence

5

,

pp. 100172

.

Cited by:

§3.1

.

I. Bråten (2010)

Personal epistemology in education: concepts, issues, and implications

.

Cited by:

§1

,

§1

,

§2.2

,

§8.2.1

.

V. Braun and V. Clarke (2019)

Reflecting on reflexive thematic analysis

.

Qualitative Research in Sport, Exercise and Health

11

(

4

),

pp. 589–597

.

Cited by:

§3.3

.

V. Braun and V. Clarke (2021)

One size fits all? what counts as quality practice in (reflexive) thematic analysis?

.

Qualitative Research in Psychology

18

(

3

),

pp. 328–352

.

Cited by:

§3.3

.

P. C. Brown, H. L. Roediger III, and M. A. McDaniel (2014)

Make it stick: the science of successful learning

.

Harvard University Press

.

Cited by:

§2.3

.

J. Brownlee, G. Boulton-Lewis, and N. Purdie (2002)

Core beliefs about knowing and peripheral beliefs about learning: developing an holistic conceptualisation of epistemological beliefs

.

Australian Journal of Educational and Developmental Psychology

2

,

pp. 1–16

.

Cited by:

§1

.

K. R. Butcher and T. Sumner (2011)

Self-directed learning and the sensemaking paradox

.

Human–Computer Interaction

26

(

1-2

),

pp. 123–159

.

Cited by:

§8.3.2

.

N. Carter (2014)

The use of triangulation in qualitative research

.

Number 5/September 2014

41

(

5

),

pp. 545–547

.

Cited by:

§1

.

G. Cassinadri (2024)

ChatGPT and the technology-education tension: applying contextual virtue epistemology to a cognitive artifact

.

Philosophy & Technology

37

(

1

),

pp. 1–28

.

Cited by:

§2.2

.

W. Charlton

et al.

(1983)

Aristotle’s physics: books i and ii

.

Oxford University Press

.

Cited by:

§5.2.1

.

S. Chavanayarn (2023)

Navigating ethical complexities through epistemological analysis of chatgpt

.

Bulletin of Science, Technology & Society

43

(

3-4

),

pp. 105–114

.

Cited by:

§2.2

.

J. Chen, X. Lu, Y. Du, M. Rejtig, R. Bagley, M. Horn, and U. Wilensky (2024a)

Learning agent-based modeling with llm companions: experiences of novices and experts using chatgpt & netlogo chat

.

In

Proceedings of the CHI Conference on Human Factors in Computing Systems

,

pp. 1–18

.

Cited by:

§2.3

.

W. Chen, C. Yu, H. Wang, Z. Wang, L. Yang, Y. Wang, W. Shi, and Y. Shi (2023a)

From gap to synergy: enhancing contextual understanding through human-machine collaboration in personalized systems

.

In

Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology

,

pp. 1–15

.

Cited by:

§1

.

Y. Chen, S. Jensen, L. J. Albert, S. Gupta, and T. Lee (2023b)

Artificial intelligence (ai) student assistants in the classroom: designing chatbots to support student success

.

Information Systems Frontiers

25

(

1

),

pp. 161–182

.

Cited by:

§1

.

Z. Chen, J. Wang, M. Xia, K. Shigyo, D. Liu, R. Zhang, and H. Qu (2024b)

StuGPTViz: a visual analytics approach to understand student-chatgpt interactions

.

arXiv preprint arXiv:2407.12423

.

Cited by:

§1

,

§8.2.1

.

M. T. Chi, J. Adams, E. B. Bogusch, C. Bruchok, S. Kang, M. Lancaster, R. Levy, N. Li, K. L. McEldoon, G. S. Stump,

et al.

(2018)

Translating the icap theory of cognitive engagement into practice

.

Cognitive Science

42

(

6

),

pp. 1777–1832

.

Cited by:

§2.3

,

§8.2.2

.

M. T. Chi and R. Wylie (2014)

The icap framework: linking cognitive engagement to active learning outcomes

.

Educational psychologist

49

(

4

),

pp. 219–243

.

Cited by:

§2.3

,

§2.3

,

§8.1

.

M. T. Chi (2009)

Active-constructive-interactive: a conceptual framework for differentiating learning activities

.

Topics in Cognitive Science

1

(

1

),

pp. 73–105

.

Cited by:

§2.3

,

§8.1

.

V. Clarke and V. Braun (2017)

Thematic analysis

.

The Journal of Positive Psychology

12

(

3

),

pp. 297–298

.

Cited by:

§3.1

.

S. M. Cohen and C. D. Reeve (2000)

Aristotle’s metaphysics

.

Cited by:

§5.2.1

.

C. Cohn, C. Snyder, J. Montenegro, and G. Biswas (2024)

Towards a human-in-the-loop llm approach to collaborative discourse analysis

.

In

International Conference on Artificial Intelligence in Education

,

pp. 11–19

.

Cited by:

§2.3

.

J. Creighton (2023)

External Links:

Link

Cited by:

§3.2.2

.

V. Cucuiat and J. Waite (2024)

Feedback literacy: holistic analysis of secondary educators’ views of llm explanations of program error messages

.

In

Proceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1

,

pp. 192–198

.

Cited by:

item

C1

.

.

M. Demirbag (2021)

Modeling the relations among argumentativeness, epistemological beliefs and self-regulation skills.

.

International Journal of Progressive Education

17

(

4

),

pp. 327–340

.

Cited by:

§2.1

.

Y. Deng, W. Lei, M. Huang, and T. Chua (2023)

Rethinking conversational agents in the era of llms: proactivity, non-collaborativity, and beyond

.

In

Proceedings of the Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region

,

New York, NY, USA

.

External Links:

Link

Cited by:

§2.1

.

R. A. Duschl, R. Hamilton, and R. E. Grandy (1990)

Psychology and epistemology: match or mismatch when applied to science education?

.

International Journal of Science Education

12

(

3

),

pp. 230–243

.

Cited by:

§5.2.1

.

A. Falcon (2006)

Aristotle on causality

.

Cited by:

§1

,

§5.2.1

.

A. J. Fiannaca, C. Kulkarni, C. J. Cai, and M. Terry (2023)

Programming without a programming language: challenges and opportunities for designing developer tools for prompt programming

.

In

Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems

,

pp. 1–7

.

Cited by:

§2.1

,

§2.3

,

§8.1

,

§8

.

N. J. Fratiwi, A. Samsudin, T. R. Ramalis, A. Saregar, R. Diani, K. Ravanis,

et al.

(2020)

Developing memori on newton’s laws: for identifying students’ mental models.

.

European Journal of Educational Research

9

(

2

),

pp. 699–708

.

Cited by:

§5.2.1

.

Y. Fu, Z. Weng, and J. Wang (2024)

Examining ai use in educational contexts: a scoping meta-review and bibliometric analysis

.

International Journal of Artificial Intelligence in Education

,

pp. 1–57

.

Cited by:

§2.1

.

P. Galdas (2017)

Revisiting bias in qualitative research: reflections on its relationship with funding and impact

.

Vol.

16

,

SAGE Publications Sage CA: Los Angeles, CA

.

Cited by:

§8.3.1

.

W. Gan, Z. Qi, J. Wu, and J. C. Lin (2023)

LLMs in education: vision and opportunities

.

In

2023 IEEE international conference on big data (BigData)

,

pp. 4776–4785

.

Cited by:

§2.1

.

L. Gao, J. Lu, Z. Shao, Z. Lin, S. Yue, C. Ieong, Y. Sun, R. J. Zauner, Z. Wei, and S. Chen (2024)

Fine-tuned large language model for visualization system: a study on self-regulated learning in education

.

arXiv preprint arXiv:2407.20570

.

Cited by:

§1

,

§2.1

,

§8.2.2

,

§8

.

K. I. Gero, C. Swoopes, Z. Gu, J. K. Kummerfeld, and E. L. Glassman (2024)

Supporting sensemaking of large language model outputs at scale

.

In

Proceedings of the CHI Conference on Human Factors in Computing Systems

,

pp. 1–21

.

Cited by:

§1

,

§2.2

.

E. L. Glassman, J. Scott, R. Singh, P. J. Guo, and R. C. Miller (2015)

OverCode: visualizing variation in student solutions to programming problems at scale

.

ACM Transactions on Computer-Human Interaction (TOCHI)

22

(

2

),

pp. 1–35

.

Cited by:

§2.3

.

A. Goslen, Y. J. Kim, J. Rowe, and J. Lester (2024)

LLM-based student plan generation for adaptive scaffolding in game-based learning environments

.

International Journal of Artificial Intelligence in Education

,

pp. 1–26

.

Cited by:

§2.1

.

E. Groenland (2018)

Employing the matrix method as a tool for the analysis of qualitative research data in the business domain

.

International Journal of Business and Globalisation

21

(

1

),

pp. 119–134

.

Cited by:

§4.2.1

.

G. Guest, K. M. MacQueen, and E. E. Namey (2011)

Applied thematic analysis

.

sage publications

.

Cited by:

§3.3

.

M. Guo, Z. Zhou, D. Gotz, and Y. Wang (2023)

Grafs: graphical faceted search system to support conceptual understanding in exploratory search

.

ACM Transactions on Interactive Intelligent Systems

13

(

2

),

pp. 1–36

.

Cited by:

§6.2

.

R. Hadi Mogavi, Y. Zhang, E. Haq, Y. Wu, P. Hui, and X. Ma (2022)

What do users think of promotional gamification schemes? a qualitative case study in a question answering website

.

Proceedings of the ACM on Human-Computer Interaction

6

(

CSCW2

),

pp. 1–34

.

Cited by:

§8.3.1

.

K. Hammarberg, M. Kirkman, and S. de Lacey (2016)

Qualitative research methods: when to use them and how to judge them

.

Human reproduction

31

(

3

),

pp. 498–501

.

Cited by:

§8.3.1

.

K. Hartley, M. Hayak, and U. H. Ko (2024)

Artificial intelligence supporting independent student learning: an evaluative case study of chatgpt and learning to code

.

Education Sciences

14

(

2

),

pp. 120

.

Cited by:

§2.1

.

H. A. He, J. Walny, S. Thoma, S. Carpendale, and W. Willett (2023)

Enthusiastic and grounded, avoidant and cautious: understanding public receptivity to data and visualizations

.

IEEE Transactions on Visualization and Computer Graphics

.

Cited by:

§3.3.3

.

M. Hocutt (1974)

Aristotle’s four becauses

.

Philosophy

49

(

190

),

pp. 385–399

.

Cited by:

§1

,

§5.2.1

.

B. K. Hofer and P. R. Pintrich (1997)

The development of epistemological theories: beliefs about knowledge and knowing and their relation to learning

.

Review of Educational Research

67

(

1

),

pp. 88–140

.

Cited by:

§1

,

§2.2

,

§4.2.1

.

J. Hu, J. Guo, N. Tang, X. Ma, Y. Yao, C. Yang, and Y. Xu (2024)

Designing the conversational agent: asking follow-up questions for information elicitation

.

Proceedings of the ACM on Human-Computer Interaction

8

(

CSCW1

),

pp. 1–30

.

Cited by:

§2.1

,

§8.2.2

,

§8

.

C. L. Huang, C. Wu, and S. C. Yang (2023)

How students view online knowledge: epistemic beliefs, self-regulated learning and academic misconduct

.

Computers & Education

200

,

pp. 104796

.

Cited by:

§2.1

,

§3.3.2

,

§3.3

,

§4.2.1

.

W. Huang, K. F. Hew, and L. K. Fryer (2022)

Chatbots for language learning—are they really useful? a systematic review of chatbot-supported language learning

.

Journal of Computer Assisted Learning

38

(

1

),

pp. 237–257

.

Cited by:

§1

.

S. E. Huber, K. Kiili, S. Nebel, R. M. Ryan, M. Sailer, and M. Ninaus (2024)

Leveraging the potential of llms in education through playful and game-based learning

.

Educational Psychology Review

36

(

1

),

pp. 25

.

Cited by:

§2.1

.

H. Jin, S. Lee, H. Shin, and J. Kim (2024)

Teach ai how to code: using large language models as teachable agents for programming education

.

In

Proceedings of the CHI Conference on Human Factors in Computing Systems

,

pp. 1–28

.

Cited by:

§2.2

.

S. A. Karabenick and M. H. Dembo (2011)

Understanding and facilitating self-regulated help seeking

.

New Directions for Teaching and Learning

2011

(

126

),

pp. 33–43

.

Cited by:

§2.3

.

S. Kim (2023)

Using chatgpt for language editing in scientific articles

.

Maxillofacial plastic and reconstructive surgery

45

(

1

),

pp. 13

.

Cited by:

§1

.

B. Kotzee (2018)

Applied epistemology of education

.

In

The Routledge Handbook of Applied Epistemology

,

pp. 211–230

.

Cited by:

§5.2.1

.

S. Krasmann (2020)

The logic of the surface: on the epistemology of algorithms in times of big data

.

Information, Communication & Society

23

(

14

),

pp. 2096–2109

.

Cited by:

§1

,

§8.2.1

.

H. Kumar, I. Musabirov, M. Reza, J. Shi, A. Kuzminykh, J. J. Williams, and M. Liut (2023)

Impact of guidance and interaction strategies for llm use on learner performance and perception

.

arXiv preprint arXiv:2310.13712

.

Cited by:

§2.1

,

§8.2.2

.

A. Lacerda, S. A. A. Freitas, and C. S. Ramos (2024)

Gamified chatbot management process: a way to build gamified chatbots

.

In

Intelligent Systems Conference

,

pp. 18–36

.

Cited by:

§8.2.2

.

S. Y. Lee and K. Ma (2024)

HINTs: sensemaking on large collections of documents with hypergraph visualization and intelligent agents

.

arXiv preprint arXiv:2403.02752

.

Cited by:

§2.2

.

U. Lee, H. Jung, Y. Jeon, Y. Sohn, W. Hwang, J. Moon, and H. Kim (2024)

Few-shot is enough: exploring chatgpt prompt engineering method for automatic question feneration in english education

.

Education and Information Technologies

29

(

9

),

pp. 11483–11515

.

Cited by:

§2.1

.

H. Li, T. Xu, C. Zhang, E. Chen, J. Liang, X. Fan, H. Li, J. Tang, and Q. Wen (2024a)

Bringing generative ai to adaptive learning in education

.

arXiv preprint arXiv:2402.14601

.

Cited by:

§2.1

.

Z. Li, F. Li, Q. Fu, X. Wang, H. Liu, Y. Zhao, and W. Ren (2024b)

LLMs and medical education: a paradigm shift in educator roles

.

Smart Learning Environments

11

(

1

),

pp. 26

.

Cited by:

§2.1

.

Z. Li, M. Liang, H. T. Le, R. Lc, and Y. Luo (2023)

Exploring design opportunities for reflective conversational agents to reduce compulsive smartphone use

.

In

Proceedings of the 5th International Conference on Conversational User Interfaces

,

New York, NY, USA

,

pp. 1–6

.

External Links:

Link

Cited by:

§2.1

.

P. Limna, T. Kraiwanit, K. Jangjarat, P. Klayklung, and P. Chocksathaporn (2023)

The use of chatgpt in the digital era: perspectives on chatbot implementation

.

Journal of Applied Learning and Teaching

6

(

1

).

Cited by:

§2.1

.

X. Lin (2023)

Exploring the role of chatgpt as a facilitator for motivating self-directed learning among adult learners

.

Adult Learning

,

pp. 10451595231184928

.

Cited by:

§2.1

.

M. X. Liu, T. Wu, T. Chen, F. M. Li, A. Kittur, and B. A. Myers (2024a)

Selenite: scaffolding online sensemaking with comprehensive overviews elicited from large language models

.

In

Proceedings of the CHI Conference on Human Factors in Computing Systems

,

pp. 1–26

.

Cited by:

§1

.

Y. Liu, P. Sharma, M. J. Oswal, H. Xia, and Y. Huang (2024b)

Personaflow: boosting research ideation with llm-simulated expert personas

.

arXiv preprint arXiv:2409.12538

.

Cited by:

§2.1

,

§8.2.2

.

Z. Liu, S. X. Yin, C. Lee, and N. F. Chen (2024c)

Scaffolding language learning via multi-modal tutoring systems with pedagogical instructions

.

arXiv preprint arXiv:2404.03429

.

Cited by:

§2.1

.

B. Luo, R. Y. Lau, C. Li, and Y. Si (2022)

A critical review of state-of-the-art chatbot designs and applications

.

Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery

12

(

1

),

pp. e1434

.

Cited by:

§1

.

X. Ma, S. Mishra, A. Liu, S. Y. Su, J. Chen, C. Kulkarni, H. Cheng, Q. Le, and E. Chi (2024)

Beyond chatbots: explorellm for structured thoughts and personalized model responses

.

In

Extended Abstracts of the CHI Conference on Human Factors in Computing Systems

,

pp. 1–12

.

Cited by:

§2.2

.

J. Macallister (2012)

Virtue epistemology and the philosophy of education

.

Journal of Philosophy of Education

46

(

2

),

pp. 251–270

.

Cited by:

§5.2.1

.

G. Marchionini (2019)

Search, sensemaking and learning: closing gaps

.

Information and Learning Sciences

120

(

1/2

),

pp. 74–86

.

Cited by:

§2.2

.

R. H. Mogavi, C. Deng, J. J. Kim, P. Zhou, Y. D. Kwon, A. H. S. Metwally, A. Tlili, S. Bassanelli, A. Bucchiarone, S. Gujar,

et al.

(2024)

ChatGPT in education: a blessing or a curse? a qualitative study exploring early adopters’ utilization and perceptions

.

Computers in Human Behavior: Artificial Humans

2

(

1

),

pp. 100027

.

Cited by:

§2.3

,

§8.1

.

H. Mokhtari (2014)

A quantitative survey on the influence of students’ epistemic beliefs on their general information seeking behavior

.

The Journal of Academic Librarianship

40

(

3-4

),

pp. 259–263

.

Cited by:

§3.3.2

,

§3.3

,

§4.2.1

,

§8.2.1

.

T. T. Moores and J. C. Chang (2009)

Self-efficacy, overconfidence, and the negative effect on subsequent performance: a field study

.

Information & Management

46

(

2

),

pp. 69–76

.

Cited by:

§2.3

.

C. H. Pérez and T. Ziemke (2007)

Aristotle, autonomy and the explanation of behaviour

.

Pragmatics & cognition

15

(

3

),

pp. 547–571

.

Cited by:

§5.2.1

.

M. Peters (2000)

Does constructivist epistemology have a place in nurse education?

.

Journal of Nursing Education

39

(

4

),

pp. 166–172

.

Cited by:

§1

,

§8.2.1

.

P. Pirolli and S. Card (2005)

The sensemaking process and leverage points for analyst technology as identified through cognitive task analysis

.

In

Proceedings of International Conference on Intelligence Analysis

,

Vol.

5

,

pp. 2–4

.

Cited by:

§2.2

,

§6.3.2

,

§8.2.3

,

§8.3.2

.

K. M. Reddy and R. Guha (2023)

Automatic text summarization for conversational chatbot

.

In

2023 IEEE 8th International Conference for Convergence in Technology (I2CT)

,

pp. 1–7

.

Cited by:

§1

.

E. Z. Rong, M. M. Zhou, G. Gao, and Z. Lu (2023)

Understanding personal data tracking and sensemaking practices for self-directed learning in non-classroom and non-computer-based contexts

.

In

Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems

,

pp. 1–16

.

Cited by:

§2.2

.

C. Rosé, Y. Wang, Y. Cui, J. Arguello, K. Stegmann, A. Weinberger, and F. Fischer (2008)

Analyzing collaborative learning processes automatically: exploiting the advances of computational linguistics in computer-supported collaborative learning

.

International Journal of Computer-supported Collaborative Learning

3

,

pp. 237–271

.

Cited by:

§2.2

.

C. W. Safranek, A. E. Sidamon-Eristoff, A. Gilson, and D. Chartash (2023)

The role of llm in medical education: applications and implications

.

Vol.

9

,

JMIR Publications Toronto, Canada

.

Cited by:

§2.1

.

M. Sailer, R. Maier, S. Berger, T. Kastorff, and K. Stegmann (2024)

Learning activities in technology-enhanced learning: a systematic review of meta-analyses and second-order meta-analysis in higher education

.

Learning and Individual Differences

112

,

pp. 102446

.

Cited by:

§2.3

.

M. Sailer, F. Schultz-Pernice, and F. Fischer (2021)

Contextual facilitators for learning activities involving technology in higher education: the c

♭

\flat

-model

.

Computers in Human Behavior

121

,

pp. 106794

.

Cited by:

§2.3

.

P. Schneider, A. Afzal, J. Vladika, D. Braun, and F. Matthes (2023)

Investigating conversational search behavior for domain exploration

.

In

European Conference on Information Retrieval

,

pp. 608–616

.

Cited by:

§3.3

.

M. Schommer-Aikins (2004)

Explaining the epistemological belief system: introducing the embedded systemic model and coordinated research approach

.

Educational Psychologist

39

(

1

),

pp. 19–29

.

Cited by:

§2.2

.

C. Schulte and L. Budde (2018)

A framework for computing education: hybrid interaction system: the need for a bigger picture in computing education

.

In

Proceedings of the 18th Koli Calling International Conference on Computing Education Research

,

pp. 1–10

.

Cited by:

§2.1

,

§2.3

.

J. Schulze (2024)

External Links:

Link

Cited by:

§6.2

.

C. Shah and E. M. Bender (2022)

Situating search

.

In

Proceedings of the 2022 Conference on Human Information Interaction and Retrieval

,

pp. 221–232

.

Cited by:

§2.1

,

§8.1

,

§8

.

S. Shankar, J. Zamfirescu-Pereira, B. Hartmann, A. G. Parameswaran, and I. Arawjo (2024)

Who validates the validators? aligning llm-assisted evaluation of llm outputs with human preferences

.

arXiv preprint arXiv:2404.12272

.

Cited by:

§1

.

R. Sheng, L. Yang, H. Li, Y. Luo, Z. Xu, Z. Zhou, D. Gotz, and H. Qu (2023)

Knowledge compass: a question answering system guiding students with follow-up question recommendations

.

In

Adjunct Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology

,

pp. 1–4

.

Cited by:

§1

,

§8.2.1

.

S. R. Shihab, N. Sultana, and A. Samad (2023)

Revisiting the use of chatgpt in business and educational fields: possibilities and challenges

.

BULLET: Jurnal Multidisiplin Ilmu

2

(

3

),

pp. 534–545

.

Cited by:

§1

.

S. Shirzad, H. Barjesteh, M. Dehqan, and M. Zare (2022)

Epistemic beliefs and learners’ self-efficacy as predictors of language learning strategies: toward testing a model

.

Frontiers in Psychology

13

,

pp. 867560

.

Cited by:

§2.2

.

C. S. Sibilin (2023)

Education and the epistemological crisis in the age of chatgpt

.

Critical Review

,

pp. 1–12

.

Cited by:

§2.2

.

T. L. Smestad and F. Volden (2019)

Chatbot personalities matters: improving the user experience of chatbot interfaces

.

In

Internet Science: INSCI 2018 International Workshops, St. Petersburg, Russia, October 24–26, 2018, Revised Selected Papers 5

,

pp. 170–181

.

Cited by:

§2.3

.

StudySmarter (2024)

External Links:

Link

Cited by:

§6.2

.

S. Suh, B. Min, S. Palani, and H. Xia (2023)

Sensecape: enabling multilevel exploration and sensemaking with llm

.

In

Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology

,

pp. 1–18

.

Cited by:

§1

,

§1

,

§2.2

,

§6.2

,

§8.2.3

,

§8

.

I. Taylor, M. Barker, and A. Jones (2003)

Promoting mental model building in astronomy education

.

International Journal of Science Education

25

(

10

),

pp. 1205–1225

.

Cited by:

§5.2.1

.

J. E. Taylor and E. Sondermeyer (2023)

Using aristotle’s four causes to evaluate and revise adult education programs

.

Adult Learning

34

(

4

),

pp. 244–255

.

Cited by:

§5.2.1

.

K. Urgo and J. Arguello (2022)

Learning assessments in search-as-learning: a survey of prior work and opportunities for future research

.

Information Processing & Management

59

(

2

),

pp. 102821

.

Cited by:

§2.3

.

V. Venkatesh and H. Bala (2008)

Technology acceptance model 3 and a research agenda on interventions

.

Decision sciences

39

(

2

),

pp. 273–315

.

Cited by:

§6.3.3

.

K. Vredenburg, J. Mao, P. W. Smith, and T. Carey (2002)

A survey of user-centered design practice

.

In

Proceedings of the SIGCHI conference on Human factors in computing systems

,

pp. 471–478

.

Cited by:

§3

.

X. Wang, D. Yang, M. Wen, K. Koedinger, and C. P. Rosé (2015)

Investigating how student’s cognitive behavior in mooc discussion forums affect learning gains

.

International educational data mining society

.

Cited by:

§2.1

,

§2.3

.

C. Wekerle, M. Daumiller, and I. Kollar (2022)

Using digital technology to promote higher education learning: the importance of different learning activities and their relations to learning outcomes

.

Journal of Research on Technology in Education

54

(

1

),

pp. 1–17

.

Cited by:

§2.1

,

§2.3

.

J. White, Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. Elnashar, J. Spencer-Smith, and D. C. Schmidt (2023)

A prompt pattern catalog to enhance prompt engineering with chatgpt

.

arXiv preprint arXiv:2302.11382

.

Cited by:

§5.2.2

.

B. P. Woolf, H. C. Lane, V. K. Chaudhri, and J. L. Kolodner (2013)

AI grand challenges for education

.

AI magazine

34

(

4

),

pp. 66–84

.

Cited by:

§2.2

.

J. Zamfirescu-Pereira, R. Y. Wong, B. Hartmann, and Q. Yang (2023)

Why johnny can’t prompt: how non-ai experts try (and fail) to design llm prompts

.

In

Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems

,

pp. 1–21

.

Cited by:

§2.1

,

§2.3

,

§2.3

,

item

C2

.

,

§8.1

,

§8

.

P. Zhang and D. Soergel (2020)

Cognitive mechanisms in sensemaking: a qualitative user study

.

Journal of the Association for Information Science and Technology

71

(

2

),

pp. 158–171

.

Cited by:

§2.2

.

Z. Zhang, J. Gao, R. S. Dhaliwal, and T. J. Li (2023)

Visar: a human-ai argumentative writing assistant with visual programming and rapid draft prototyping

.

In

Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology

,

pp. 1–30

.

Cited by:

§1

,

§2.2

.

C. Zheng, K. Yuan, B. Guo, R. H. Mogavi, Z. Peng, S. Ma, and X. Ma (2024a)

Charting the future of ai in project-based learning: a co-design exploration with students

.

arXiv preprint arXiv:2401.14915

.

Cited by:

§1

.

C. Zheng, Y. Zhang, Z. Huang, C. Shi, M. Xu, and X. Ma (2024b)

DiscipLink: unfolding interdisciplinary information seeking process via human-ai co-exploration

.

In

Proceedings of the CHI Conference on Human Factors in Computing Systems

,

pp. 1–15

.

Cited by:

§1

,

§1

,

§2.1

,

§2.2

,

§8.2.3

,

§8

.

D. Zollman, A. Sirnoorkar, and J. T. Laverty (2023)

Analyzing ai and student responses through the lens of sensemaking and mechanistic reasoning

.

In

Proceedings of the Physics Education Research Conference (PERC)

,

pp. 415–420

.

Cited by:

§5.2.1

.

Appendix A

Formative Study

A.1.

Demographics of Semi-structured Interview Participants

Table 6.

The table presents demographic information (age, gender, education) and expertise of formative study participants, along with their years of academic research (YAR), familiarity with large language models (LLM), time spent on self-learning tasks (SLT), quiz time (Quiz T), and quiz scores.

ID

Gender

Age

Edu. Background

Expertise

YAR

LLM Familiarity

SLT (mins)

Quiz T (mins)

Quiz Score

P1

Female

29

Doctorate degree

Social Science

5+

Somewhat Familiar

23

10

7/10

P2

Female

26

Bachelor’s degree

STEM

3-5

Somewhat Familiar

40

11

7/10

P3

Female

29

Master’s degree

Social Science

≤

3

\leq 3

Somewhat Familiar

21

9

8/10

P4

Female

23

Master’s degree

Interdisciplinary Studies

≤

3

\leq 3

Very Familiar

20

9

7.5/10

P5

Male

26

Doctorate degree

STEM

3-5

Somewhat Familiar

18.5

7

8/10

P6

Female

29

Master’s degree

Social Science

≤

3

\leq 3

Somewhat Familiar

45

15

7.5/10

P7

Male

29

Master’s degree

STEM

3-5

Very Unfamiliar

18

10

7/10

P8

Male

29

Master’s degree

STEM

N/A

Very Familiar

18

8

7/10

P9

Male

25

Master’s degree

Interdisciplinary Studies

≤

3

\leq 3

Very Familiar

30

8

9/10

P10

Male

25

Bachelor’s degree

STEM

3-5

Extremely Familiar

16

4

7/10

P11

Female

27

Master’s degree

Interdisciplinary Studies

N/A

Somewhat Unfamiliar

17

4

7/10

P12

Male

32

Doctorate degree

STEM

5 +

Very Familiar

17

7

8.5/10

P13

Male

27

Doctorate degree

STEM

3-5

Very Familiar

23

8

5.5/10

P14

Male

28

Doctorate degree

Medical

3-5

Somewhat Unfamiliar

24

7

5/10

P15

Female

27

Master’s degree

Social Science

3-5

Somewhat Familiar

26

9

5/10

P16

Female

26

Master’s degree

STEM

≤

3

\leq 3

Somewhat Familiar

20

6

5.5/10

P17

Male

25

Master’s degree

STEM

≤

3

\leq 3

Very Familiar

18

8

8/10

P18

Female

29

Bachelor’s degree

Medical

5 +

Very Unfamiliar

20

15

5/10

P19

Male

24

Master’s degree

Interdisciplinary Studies

≤

3

\leq 3

Extremely Familiar

20

6

8.5/10

P20

Female

27

Master’s degree

Social Science

3-5

Somewhat Familiar

16

5

8.5/10

P21

Male

28

Master’s degree

STEM

5+

Somewhat Familiar

15

9

5/10

P22

Female

24

Master’s degree

STEM

≤

3

\leq 3

Extremely Familiar

17.5

7

6/10

P23

Female

32

Doctorate degree

STEM

5+

Somewhat Familiar

23

4

9/10

P24

Female

28

Master’s degree

Social Science

≤

3

\leq 3

Somewhat Familiar

20

5

6.5/10

P25

Female

24

Bachelor’s degree

STEM

≤

3

\leq 3

Extremely Familiar

19

7

5.5/10

P26

Female

27

Master’s degree

Social Science

3-5

Somewhat Unfamiliar

42

6

5/10

A.2.

Web Page for the Self-learning Task

Figure 10.

Design of the probe system for self-learning tasks in the formative study. This includes 1) reading materials; 2) a table of contents for navigating different subsections; 3) a question-and-answer chatbot; 4) a

Save

button for downloading dialogue records; and (5) a

Start the Quiz

button to exit self-learning tasks and begin the quiz. The

Save

button and the

Start the Quiz

button become active 15 minutes after entering the web page.

Appendix B

Comparing Queries Across Two Conditions

Table 7.

This table presents follow-up questions on two topics, NFTs and semiotics, generated by

CausaDisco

and GPT-4o (baseline) during user evaluation. Questions in bold represent those selected by users in multi-turn conversations.

Seed Query by Users

Condition

Generated Follow-Up Questions by LLM

How do creators of

NFTs determine the

prices of their products?

CausaDisco

C-1. How do market trends influence the initial pricing strategy of an

NFT creator?

C-2. What impact does the perceived future value have on setting the price of

a new NFT?

C-3. In what ways do the creator’s reputation and past sales history affect the

pricing of their NFTs?

C-4. How does the utility or functionality of an NFT contribute to its valuation

by the creator?

Baseline

B-1. What are some common platforms for buying and selling NFTs?

B-2. How do creators market their NFT to attract buyers?

B-3. What role does social media play in NFT pricing?

CausaDisco

C-1-1. How do fluctuations in cryptocurrency value affect the pricing

strategies of NFT creators?

C-1-2. What role do collector and investor behaviors play in shaping market

trends for NFTs?

C-1-3. How can NFT creators leverage social media and community

engagement to influence market trends and pricing?

C-1-4. In what ways do historical sales data and analytics tools aid NFT

creators in understanding and adapting to market trends?

Baseline

B-3-1. How do different social media platforms impact NFT pricing differently?

B-3-2. What strategies do creators use to leverage social media for pricing

their NFTs?

B-3-3. Can you explain how market trends on social media are identified

and analyzed?

What is semiotics in

concise summary?

CausaDisco

C-1. What are the key components of a sign in semiotic analysis?

C-2. How does an individual’s cultural context shape their understanding

of specific signs?

C-2. What role do technological changes play in altering the significance and

use of signs in communication?

C-4. How can semiotic knowledge enhance the effectiveness of communication

across diverse media formats?

Baseline

B-1. How does semiotics apply to everyday communication?

B-2. What are some key theories or figures in semiotics?

B-3. Can you explain the difference between denotation and connotation?

CausaDisco

C-1-1. How does the relationship between the signifier and the

signified influence the interpretation of a sign?

C-1-2. In what ways can the context alter the perceived relationship

between the signifier and the signified?

C-1-3. How do social and cultural factors affect the creation and

interpretation of signs?

C-1-4. Can the evolution of language and symbols disrupt the traditional

signifier-signified relationship?

Baseline

B-2-1. How do these theories apply to modern digital communication?

B-2-2. Can you explain Barthes’ concept of myth in more detail?

B-2-3. What are some practical applications of semiotics in advertising?

BETA