Title: 2412.15660
ArXiv: 2412.15660

Adaptable and Precise: Enterprise-Scenario LLM Function-Calling Capability Training Pipeline

\addbibresource

ref.bib

Adaptable and Precise: Enterprise-Scenario LLM Function-Calling Capability Training Pipeline

Guancheng Zeng

1

Digital China AI Research

{cenggc

1

, huhwa

2

}@digitalchina.com

Wentao Ding

Digital China AI Research

{cenggc

1

, huhwa

2

}@digitalchina.com

Beining Xu

Digital China AI Research

{cenggc

1

, huhwa

2

}@digitalchina.com

Chi Zhang

Digital China AI Research

{cenggc

1

, huhwa

2

}@digitalchina.com

Wenqiang Han

Digital China AI Research

{cenggc

1

, huhwa

2

}@digitalchina.com

Gang Li

Digital China AI Research

{cenggc

1

, huhwa

2

}@digitalchina.com

Jingjing Mo

Digital China AI Research

{cenggc

1

, huhwa

2

}@digitalchina.com

Pengxu Qiu

Digital China AI Research

{cenggc

1

, huhwa

2

}@digitalchina.com

Xinran Tao

Digital China AI Research

{cenggc

1

, huhwa

2

}@digitalchina.com

Wang Tao

Digital China AI Research

{cenggc

1

, huhwa

2

}@digitalchina.com

Haowen Hu

2†

†Corresponding Author

Digital China AI Research

{cenggc

1

, huhwa

2

}@digitalchina.com

Abstract

Enterprises possess a vast array of API assets scattered across various functions, forming the backbone of existing business processes. By leveraging these APIs as functional tools, enterprises can design diverse, scenario-specific agent applications, driven by on-premise function-calling models as the core engine. However, generic models often fail to meet enterprise requirements in terms of computational efficiency, output accuracy, and stability, necessitating scenario-specific adaptation. In this paper, we propose a training pipeline for function-calling capabilities tailored to real-world business scenarios. This pipeline includes the synthesis and augmentation of scenario-specific function-calling data, model fine-tuning, and performance evaluation and analysis. Using this pipeline, we generated 1,260 fully AI-generated samples and 1,035 augmented manually-labeled samples in digital HR agent scenario. The

Qwen2.5-Coder-7B-Instruct

model was employed as the base model and fine-tuned using the LoRA method on four GPUs with 24GB VRAM. Our fine-tuned model demonstrated outstanding performance in evaluations and practical applications, surpassing

GPT-4

and

GPT-4o

in accuracy on the test set. These results validate the reliability of the proposed pipeline for training scenario-specific function-calling models.

1

Introduction

In the era of generative AI, AI agents are defined as autonomous systems capable of handling tasks independently

[

yao2022react

,

xi2023rise

]

. These agents can not only respond directly to user queries but also perform complex task decomposition and execution by invoking various functional tools

[

wang2024tdag

,

jeyakumar2024advancing

]

. With a vast knowledge base and the ability to interact with external environments, AI agents demonstrate significant potential for deep integration with existing enterprise systems

[

5331703

]

. Across industries, enterprises are actively exploring the integration of AI agent systems into real business scenarios

[

milosevicenterprise

]

. This is driven by the abundance of API assets within enterprise systems, which, when leveraged as functional tools, enable AI agents to deeply integrate with existing systems, automate workflows across specialized business scenarios, and enhance operational efficiency.

AI agents require LLMs (Large Language Models) as their core reasoning engines to generate instructions for invoking functions and APIs

[

schwartz2023enhancing

]

. While both open-source and commercial LLMs perform well in general function-calling tasks given the extensive training data

[

chung2024scaling

,

taori2023alpaca

]

, these models often struggle to provide accurate and stable function-calling instructions in specialized enterprise scenarios. Errors in instruction structure, function selection, and parameter input still occur due to the lack of domain-specific training data. Furthermore, to prevent data leakage

[

carmody2021ai

]

, enterprises typically opt to deploy on-premise LLMs rather than utilizing LLMs served on public cloud platform. However, due to the limited computational resources and lack of experience in training LLMs, it is difficult to implement AI agent effectively for Small and Medium Enterprises. Therefore, there is an urgent need to identify an efficient training pipeline suitable for developing small-scale models in enterprise scenarios, so that SMEs can easily train and deploy their own agent models based on their needs.

In this work, we designed a specialized training pipeline for function-calling capabilities tailored to professional scenarios. This pipeline includes data synthesis and augmentation for scenario-specific function-calling task, SFT (Supervised Fine-Tuning) generic function calling models
using LoRA

[

hu2021lora

]

, and comprehensive evaluation and analysis of model performance. Specifically, we used 14 workflow sets and generated 1,260 fully AI-synthesis and 1,035 manually augmented function-calling training samples in a digital HR intelligent agent system, allowing the model to better align with actual business preferences in tool selection and parameter extraction. The

Qwen2.5-Coder-7B-Instruct

model

[

bai2023qwen

]

was used as the base model and fine-tuned using LoRA within approximately five hours on four

A10

GPUs (24GB VRAM each). The fine-tuned model exhibited exceptional performance, surpassing

GPT-4

and

GPT-4o

1

1

1

https://chatgpt.com/

in function calling output structural completeness, tool selection accuracy, and parameter input accuracy, both on the test set and in real-world usage. These results demonstrated the effectiveness of our pipeline in enhancing the function-calling capabilities of moderate-size LLMs for professional scenarios under limited computational resources.

In summary, our main contributions are as follows:

•

We developed a function-calling training pipeline, enabling data synthesis, model fine-tuning, and evaluation for various enterprise-specific tools, significantly improving the performance of generic models in professional business scenario.

•

We trained a 7B function-calling LLM for the digital HR scenarios. The model exhibited outstanding performance in both evaluations and real-world usage, surpassing GPT-4 and GPT-4o in test set accuracy.

The remainder of the paper is organized as follows: Sec.

2

refers to related works of data synthesis and model evaluation on Function Calling;
Sec.

3

describes the pipeline of our work, including initial setting, data synthesis, model training and evaluation; Sec.

4

contains our experiments as well as results; ablation study and analysis are in Sec.

5

.

2

Related Work

2.1

Function-Calling Data Synthesis

The introduction of Large Language Models (LLMs) has catalyzed a significant paradigm shift in the field of deep learning

[

zhao2023survey

]

. Their ability to produce text that rivals human quality renders them invaluable for generating synthetic data, thereby alleviating the long-standing dilemma of data quality and quantity in Natural Language Processing (NLP)

[

long2024llmsdrivensyntheticdatageneration

,

wang2022self

,

xu2023wizardlm

]

.
LLMs have also demonstrated remarkable proficiency in generating function-calling data.

Toolformer

[

schick2024toolformer

]

, for instance, enhances an LLM’s ability to identify retrievable information through function calls, by annotating examples and using augmented data for training, thus enabling the model to effectively determine tool usage.
Similarly,

ToolLLaMA

[

qin2023toolllm

]

leverages over 16,000 API instances and datasets generated by ChatGPT to enhance its handling of both simple and complex multi-tool instructions, showcasing robust generalization with novel API documentation. To further simulate real user instructions and enhance the richness and complexity of commands,

ToolACE

[

liu2024toolace

]

and

ToolAlpaca

[

tang2023toolalpaca

]

employ multiple LLM agents to simulate multi-turn interactions, generating tool usage examples. Additionally, to ensure data quality, both

ToolACE

and

APIGen

[

liu2024apigen

]

implement multiple validation processes to filter data that adhere to correct invocation formats and align the semantic consistency between instruction goals and API outputs.

While LLMs with general Function-calling capabilities are emerging in abundance, there is still no standardized process for fine-tuning Function-calling LLMs in specific professional domains. To address this, we provide a comprehensive workflow for data synthesis, model fine-tuning, and model evaluation tailored to specific business scenarios. This workflow enables LLMs to better understand user instructions and generate appropriate responses by leveraging APIs effectively.

2.2

Function-Calling Evaluation Metrics

During model training, the loss function indicates the error for each forward pass, guiding the gradient updates in back-propagation. However, relying solely on loss often fails to capture the model’s task-specific performance. Researchers therefore employ discrete evaluation metrics like code execution success, multiple choice and integer scores to assess performance. The field has developed task-specific evaluation benchmarks such as

MMLU

,

CMMLU

,

C-Eval

,

GSM8K

,

ELM

, and

OpenCompass

[

hendrycks2020measuring

,

li2023cmmlu

,

huang2024c

,

cobbe2021training

,

chan2024mle

,

2023opencompass

]

, enhancing model assessment across diverse tasks like code synthesis and commonsense reasoning.

In function-calling studies, numerous evaluation methodologies have been proposed, leading to the emergence of both open-source and proprietary benchmarks. These benchmarks are designed to assess a model’s function-calling abilities, often relying on synthetic methods to construct datasets due to the scarcity of function-calling data compared to regular types of data. For instance,

APIBench

[

patil2023gorilla

]

builds its evaluation data using APIs from platforms such as TorchHub, TensorHub, and Hugging Face to assess models’ accuracy and hallucination rates in function calls. For certain task types, there is a greater emphasis on multi-turn and multi-step function-calling.

ToolBench (ToolEval)

[

xu2023tool

]

collects tool APIs from the RapidAPIHub platform and uses ChatGPT to generate diverse instructions for utilizing these APIs. Similarly,

AgentBoard

[

ma2024agentboard

]

provides nine categories of tasks and evaluates models’ multi-turn agent interaction capabilities using

ToolOperation

and

ToolQuery

. The

BFCL

[

berkeley-function-calling-leaderboard

]

framework introduced a

Live Dataset

in its V2 version, built from real-time user-contributed function documentation and queries, to evaluate models’ ability to operate effectively in dynamic environments. Its V3 version further incorporated multi-step and multi-turn evaluation logic.

For evaluating results, multiple assessment methods have been developed.

ToolBench

employs GPT-series models to evaluate model performance based on pass rates. On the other hand, benchmarks like

APIBench

and

BFCL

utilize

Abstract Syntax Trees (AST)

to parse the structure of model-generated function calls. This approach enables a multi-faceted evaluation without actually executing the call instructions, allowing for more efficient assessments. Additionally, the

BFCL

framework offers the

Exec

method, which evaluates the results of executing the model-generated function-calling instructions, making it more suitable for multi-turn dialogue evaluation. In practical use, researchers typically choose evaluation methods based on their specific needs and available resources.

3

Methodology

3.1

Overview

In enterprise environments, characterized by relatively high labor costs, it is crucial to minimize manual involvement

[

hutter2019automated

]

. Therefore, we designed a highly automated training pipeline for function-calling capabilities in enterprise models tailored to such scenarios, as shown in Figure

1

.

This pipeline consists of three main modules:

•

Data Synthesis Module:

This module generates user questions based on function tool information and creates corresponding function-calling instructions. It enhances and filters the generated questions to further improve the quantity and quality of the data

[

wei2022chain

]

. The resulting data is then assembled and divided into fine-tuning training sets and evaluation sets

[

brown2020language

]

.

•

Model Fine-Tuning Module:

Using the training set, this module fine-tunes a generic function-calling model via

LoRA

[

hu2021lora

]

and integrates the resulting weights back into the original model

[

pfeiffer2020adapterfusion

]

.

•

Model Evaluation Module:

After training, the model undergoes AST-based evaluation to assess its performance

[

allamanis2017learning

]

. Additionally, the module performs a confusion matrix analysis of the model’s function selection distribution, identifying its precision and areas of confusion for different function tools

[

sokolova2009systematic

]

.

Figure 1

:

The Overall Training Pipeline of Enterprise-Scenario Function-Calling LLM

3.2

Initial Settings

3.2.1

Generic Model

In enterprise scenarios, computational resources are often limited, and small to medium-sized enterprises typically struggle to support the fine-tuning of large-parameter models

[

houlsby2019parameter

,

han2015deep

]

. As a result, smaller-parameter models are required. Furthermore, these models need to possess general function-calling capabilities to enable the transfer of their abilities from general scenarios to enterprise-specific contexts

[

ruder2019transfer

]

. Previous research has demonstrated that models with 7B parameters can effectively handle general function-calling tasks, achieving high accuracy levels in real-world applications

[

kaplan2020scaling

]

.

3.2.2

Scenario API Tools

Depending on the task, models can utilize various types of tools, including APIs, algorithms, code workflows, operational pipelines, and even other models. The model must be able to select these tools correctly based on their format and provide the appropriate input parameters

[

gao2020making

]

. To enable stable data synthesis and tool usage, users need to supply the following information about each function tool:

Tool Name

,

Tool Description

,

Parameter Names

,

Parameter Descriptions

,

Parameter Data Types

,

Parameter Necessity

[

wang2022self

]

. Furthermore, for parameters that require high accuracy, users can enhance the stability of tool invocation by providing specific examples and default values for the input parameters

[

brown2020language

]

. Descriptions of function tools and their parameters must be provided by the user, which should clearly convey the tool’s functionality and parameter requirements. The higher the quality of these descriptions, the better the quality of the data generated, which in turn enhances the outcomes of model training

[

li2023quantity

]

.

3.2.3

Human-Annotated Seed Data

To ensure that the data closely align with real-world scenarios, the pipeline could initially incorporate a small amount of manually annotated data as seed data for the data synthesis module. Compared to fully AI-generated data, seed data annotated by business experts is better aligned with actual question-and-answer patterns, improving the quality of the generated data. This, in turn, enhances the model’s stability and accuracy

[

snow2008cheap

]

.

During manual annotation, several key aspects must be considered:

•

Diversity

: Questions should exhibit diversity in expression, including variations in sentence structure, content, and phrasing

[

wei2019eda

]

.

•

Uniqueness

: Each data entry should be unique in terms of questions and parameters, avoiding repetition of names (e.g., people, places, departments, or projects)

[

batini2009methodologies

]

.

•

Scale

: The seed data should be of sufficient size, with its volume proportionally allocated according to the importance of the function

[

kaplan2020scaling

]

.

•

Consistency

: The internal logic of the seed data must remain consistent, with uniform output formats and standardized parameter names for the same functionality

[

wang1996beyond

]

.

By adhering to these principles, the seed data ensures a high-quality foundation for data synthesis and model training, optimizing overall performance in specific application scenarios.

3.3

Data Synthesis

In this process, questions are first generated based on the descriptions of scenario-specific function tools

[

liu2024apigen

,

wei2019eda

]

. These questions, along with a small number of manually annotated question-answer pairs, are used as seed data for data augmentation

[

batini2009methodologies

]

. The augmented data are then utilized as the training set to fine-tune the model

[

gururangan2020don

]

.

3.3.1

Generating Seed Questions Based on API Description

During model fine-tuning, it has been observed that the diversity, quality, and quantity of the dataset directly influence training outcomes

[

liu2024apigen

]

. Diversity is critical for the model’s performance on unseen tasks or instructions

[

brown2020language

,

gao2020pile

]

, quality impacts its accuracy on known tasks

[

gururangan2020don

]

, and quantity effectively prevents the model from over-fitting to specific tasks. When synthesizing function-calling training datasets, the design of seed questions plays a pivotal role in determining the model’s final performance. These seed data must cover as many task scenarios as possible while adhering to the logical boundaries of the tool’s functionality. The generated questions must not deviate from the purpose described in the tool

[

brown2020language

,

gururangan2020don

]

.

To achieve this, we designed various customization prompt templates in the pipeline, incorporating elements such as role settings, generation rules, and output formats. These templates must meet the following requirements:

•

Role Setting

: Define the identity of the user asking the question, such as a salesperson, domain expert, or engineer, ensuring that the roles are diverse, clearly defined, and logically aligned with the scenario

[

ouyang2022training

]

.

•

Generation Rules

: Specify requirements for data synthesis, such as the number of questions generated per prompt, question length, required content, content to avoid, and language of the generated questions

[

gao2020pile

,

raffel2020exploring

]

.

•

Output Format

: Standardize the output format to facilitate automated data counting and dataset assembly. Unnecessary prefixes and suffixes in the generated output should be minimized

[

gururangan2020don

,

ouyang2022training

]

.

While research indicates that small-parameter models can produce more data with the same resource consumption, potentially enhancing training outcomes

[

bansal2024smaller

,

bender2020climbing

]

, their limited instruction-following capabilities make them unsuitable for complex prompt templates required in specialized scenarios. As a result, such models often produce data with formatting issues and high redundancy. Therefore, for generating initial seed data in professional contexts, high-performance models remain essential

[

raffel2020exploring

,

bender2020climbing

]

.

3.3.2

Data Augmentation Based on Seed Questions

The pipeline, starting from seed questions, broadens scenario coverage through data augmentation. This method, compared to direct annotation, lowers resource consumption. It also proves more stable than generating all data simultaneously and reduces the chance of producing duplicates. Data augmentation, like data synthesis, depends on LMs and prompt templates. However, the diversity required in augmentation directions necessitates a broader array of templates. Overall, our pipeline employs four data augmentation strategies: replacement, rewriting, simplification, and error introduction. The definition and function of these four augmentation strategies are shown in Table

1

. Samples of augmentation prompt templates in digital HR scenario can be found in

Appendix B

.

Table 1

:

Definition and Function of Augmentation Prompt Strategies

Augmentation Strategies

Definition and Function

Replacement

Non-essential elements in the seed questions are replaced without

affecting the selection of function-calling instructions or parameter

inputs.

Rewriting

The language structure of the seed questions is rearranged, altering

the grammar and phrasing of the queries while retaining their original

meaning.

Simplification

Non-essential content and conjunctions in the seed questions are

removed, shortening the original sentence into a concise query while

maintaining key information. This simulates real-world user queries

that often consist of only essential keywords.

Error Introduction

Errors such as misspellings, repetitions, or deletions of non-critical

content are added to the seed questions to help the model adapt to

occasional errors in user input.

In practice, replacement and rewriting are the most commonly used augmentation methods, while simplification and error introduction are employed less frequently. Multiple augmentation methods can be applied to the same seed question. Note that error introduction should be based on already-augmented questions to avoid excessive similarity to the original seed questions.

3.3.3

Generating Function Calling Instructions Based on Questions

After generating the questions, function-calling instructions can be created based on these questions and included in the dataset as labels and answers. The generation of instructions primarily relies on the extraction of key parameters

[

schick2024toolformer

]

. At this stage, it is generally necessary to use high-performance models for parameter extraction to ensure consistency between the extracted function-calling parameters and the questions

[

yao2022react

]

. Once the parameters are obtained, they need to be assembled into function-calling instructions in accordance with the format required by the training and testing datasets

[

berkeley-function-calling-leaderboard

,

zheng2024llamafactory

]

. This process involves integrating the parameters with the tools relevant to the problem’s context. Consequently, all augmented questions generated from the same seed question use the same function-calling instructions.

3.3.4

Data Validation and Assembly

Since certain parameters in specific problems are generated synthetically by the model and may not accurately reflect real-world conditions, the final augmented training data must undergo a validation process

[

howard2018universal

]

. This involves checking their structural integrity and verifying the accuracy of parameter assignments.

After data validation, the remaining dataset is divided into training and validation sets while ensuring that all usage scenarios of the tools are covered

[

kohavi1995study

]

. Additionally, the data format must be adjusted to align with the requirements of the training and evaluation frameworks. Typically, each data entry should include the user’s question, the model’s tool-calling instructions, and a list of available tools, formatted in JSON. Commonly used formats include

shareGPT

,

Alpaca

, and

OpenAI

formats.

Regarding the list of available tools, it must align with the task coverage of the scenarios and provide all the required tools

[

sun2019fine

]

. The number of available tools and the length of their descriptions should be kept within a certain range to avoid exceeding the model’s context window size and to reduce the complexity of tool selection

[

vaswani2017attention

]

. During data assembly, the order of the tool list in the training data should be randomized to prevent the model from over-fitting to specific sequences of tool combinations

[

srivastava2014dropout

]

.

3.4

Model Fine-tuning

3.4.1

LoRA

To prevent over-parametrization during scenario-specific task adaptation and to save computational resources, the process employs the

LoRA

(Low-Rank Adaptation) fine-tuning approach

[

hu2021lora

]

. This reduces the number of parameters required for adaptation while effectively preserving the base model’s general capabilities. During training, the model’s context length must always exceed the length of the training data, avoiding truncation of longer data entries to prevent structural disruption, which could affect training performance

[

vaswani2017attention

]

. The number of training epochs should not be excessively short or long, as this may result in under-fitting or over-fitting on the training dataset.

We use

LLaMA-Factory

[

zheng2024llamafactory

]

as the training framework in this pipeline, which supports the adaptation of various models and provides diverse fine-tuning methods. The training is primarily implemented using

PEFT

library

2

2

2

https://huggingface.co/docs/peft/en/index

. Additionally, the

DeepSpeed-ZeRO3

[

rajbhandari2020zeromemoryoptimizationstraining

]

optimization method is applied to accelerate training. This approach distributes the optimizer, gradients, and parameters across GPUs, significantly reducing memory consumption and enabling more efficient utilization of GPU resources.

3.4.2

Merging LoRA Adapters

After completing LoRA fine-tuning, the LoRA adapters are merged to the base model to facilitate further fine-tuning tasks. Additionally, the pipeline includes a

multi-adapter merging strategy

, which encompasses various processing methods such as

linear combination

,

SVD

[

stoica2024model

]

, and

parameter concatenation

[

peft

]

. These strategies enable the merging of weight matrices from different LoRA adapters prior to integration, addressing training scenarios such as multi-task fine-tuning, and data feedback fine-tuning, supports continuous updates to the model after deployment. This is crucial for business scenario as data often follows a life-cycle with continuous replacement of updated data. Furthermore, with this multi-LoRA adapter merging approach, the influence of different training data can be encapsulated within separate LoRA adapters. By selecting specific LoRA adapters, it becomes easier to mitigate the impact of certain data cycles or tasks on the model’s performance, offering enhanced flexibility in managing the model’s behavior and adaptability.

3.5

Model Evaluation

3.5.1

AST Evaluation

To effectively verify the correctness of the function-calling instructions generated by the model, evaluation must address multiple aspects, including

structural completeness of the instructions

,

accuracy of function selection

, and

accuracy of parameter inputs

. The

AST

(Abstract Syntax Tree) evaluation method enables a step-by-step parsing of the model’s output structure, allowing the correctness of function-calling instructions to be assessed without actually executing the function tools

[

aho2007compilers

]

. This capability makes AST-based evaluation particularly useful for assessing AI-generated function-calling instructions that contain fictitious parameters, a scenario where

Exec

(Execution) methods often struggle. The Exec method requires actual execution of the model’s output instructions and evaluates correctness based on the returned results

[

chen2021evaluating

]

. However, instructions with fictitious parameters—such as nonexistent IDs, names of people, places, or companies—cannot be executed, especially for query-based API tools. In addition, since AST evaluation does not involve actual API execution, it is not constrained by API response times, enabling faster verification of model outputs

[

berkeley-function-calling-leaderboard

]

.

In our pipeline, we adopted the open-source

BFCL

evaluation framework as the primary benchmark for our practical experiments, modifying it to enhance its suitability for assessing scenario-specific function-calling ability evaluation

[

berkeley-function-calling-leaderboard

]

. This framework provides a rapid and comprehensive evaluation system, including AST-based instruction parsing methods and error analysis features. These features facilitated advanced analyses of the model’s function-calling outputs, enabling deeper insights into the causes of errors and refining the evaluation process.

3.5.2

Confusion Matrix for Tool Selection

In addition to exploring the causes for errors in function-calling instructions, tool selection can be evaluated as a multi-class classification task, given that the model selects a specific tool from a list of available tools. This evaluation employs a confusion matrix to analyze the accuracy of tool selection

[

powers2020evaluation

]

. Specifically, each function tool category is treated as an independent class, and a confusion matrix is constructed where rows represent the actual tool categories and columns represent the model’s predicted categories. This approach facilitates the computation of the number of instances where the model accurately predicted a tool for each real tool’s usage scenario

[

sokolova2009systematic

]

.

Using the confusion matrix, the usage of a target tool is categorized as a positive instance, whereas the usage of alternative tools is categorized as a negative instance. This enables the calculation of

True Positives (TP)

,

True Negatives (TN)

,

False Positives (FP)

, and

False Negatives (FN)

for each tool in its respective usage scenario. These metrics, along with their definitions, are listed in Table

2

.

Table 2

:

Value and Meanings of Tool Selection Confusion Matrix

Value

Meaning in Tool Selection

TP

Selected the target tool in the target tool scenario

TN

Selected another tool in a non-target tool scenario

FP

Incorrectly selected the target tool in a non-target tool scenario

FN

Incorrectly selected another tool in the target tool scenario

Based on these values, the following key performance metrics can be derived to evaluate the model’s tool selection capability:

•

Precision

: The proportion of correct predictions among all instances where the model predicts the use of a specific tool.

Precision

=

TP

TP

+

FP

Precision

TP

TP

FP

\text{Precision}=\frac{\text{TP}}{\text{TP}+\text{FP}}

Precision = divide start_ARG TP end_ARG start_ARG TP + FP end_ARG

(1)

•

Recall

: The proportion of correct tool selections in the total number of instances where the tool is actually used in the scenario.

Recall

=

TP

TP

+

FN

Recall

TP

TP

FN

\text{Recall}=\frac{\text{TP}}{\text{TP}+\text{FN}}

Recall = divide start_ARG TP end_ARG start_ARG TP + FN end_ARG

(2)

•

F1 Score

: The harmonic mean of precision and recall, providing a comprehensive assessment of the model’s performance in selecting a specific tool.

F

1

=

2

⋅

Precision

⋅

Recall

Precision

+

Recall

subscript

𝐹

1

⋅

2

⋅

Precision

Recall

Precision

Recall

F_{1}=2\cdot\frac{\text{Precision}\cdot\text{Recall}}{\text{Precision}+\text{%
Recall}}

italic_F start_POSTSUBSCRIPT 1 end_POSTSUBSCRIPT = 2 ⋅ divide start_ARG Precision ⋅ Recall end_ARG start_ARG Precision + Recall end_ARG

(3)

These metrics reflect the model’s capability to select tools and can be employed to compare its performance before and after fine-tuning, as well as against high-performance models.

Furthermore, conducting confusion matrix analysis with evaluation results from high-performance generic models in zero-shot setting can provide insights into the correspondence between tool descriptions and user queries, as well as the confusion degree between tools. Tools with low precision rate are more likely to interfere with the selection of other tools, while tools with low recall rate are more likely to be affected by other tools. Based on these observations, practitioners can optimize the functional scope and descriptive information of tools prior to data synthesis. This ensures that each tool has clearly differentiated functionality and that its description aligns with its logical use case within the given scenario, which helps reduce tool confusion and improves the overall accuracy of tool selection by the model.

4

Experiments

4.1

Experimental Background

In our experiment, we tested our pipeline in a digital HR intelligent agent scenario within a large corporate group, involving over 8,000 employees. Users could interact with the intelligent agent in Chinese, inquiring about information related to the company’s employees and departments. The experiment provided a total of 14 specialized workflows with distinct functionalities, each encapsulated in the format of function tools. The model needs to interpret user’s queries accurately, pass them as parameters to the workflows. Workflow will automatically query the relevant data from the database, summarize the results, and then deliver them back to the users.

4.2

Experimental Setup

4.2.1

Data Synthesis

In this experiment, the training data were synthesized in two stages: seed question generation and data augmentation. These tasks were completed using GPT-4 equipped with appropriate prompt templates, aiming to maximize the quality of the generated data. Details of the prompt templates can be found in

Appendix A

and

Appendix B

.

During the data synthesis process, for each of 14 function tools (scenario workflows), 10 seed questions were generated, leads to 140 seed questions in total. These seed questions underwent data augmentation, each produce 10 enhanced question, which gave us 1400 augmented data instances. From these augmented data, we finally selected 1260 data instances as training data, 90 instances for each tools.

Furthermore, we incorporated 207 manually annotated data instances to compare the performance improvements between two types of seed data. From these 207 manually annotated seeds, we generated a total of 1,035 augmented training data instances, with 5 augmentations for each seed instance. Notably, to align with real-world scenarios usage, the data distribution of the 207 manually annotated data instances across the 14 workflows was uneven. The distribution details are shown in Figure

2

.

Figure 2

:

Data Volume and Distribution of 14 Digital HR Scenario Tools

Among the workflows,

[10]Staff Basic Information Inquiry

accounts for more than one-third of the total data volume, as it is primarily used for querying basic user information. Additionally, some workflows share similar functions and descriptions, such as

[12]Marketing Staff Achievement Inquiry

and

[8]Marketing Employee Data Inquiry

, as well as

[4]Staff Capability Analysis

,

[14]Person Job Matching Analysis

, and

[10]Staff Basic Information Inquiry

. These similarities increase the difficulty for the model to select the correct tool during function calls, leading to potential errors in tool selection. As shown in Figure

3

, both Qwen (b) and GPT-4 (c) encounter tool selection errors in certain workflows under zero-shot conditions, highlighting the complexity of this task and the necessity of fine-tuning.

Regarding data formatting, the training data was organized in the ShareGPT format. During assembly, order of the tool list was randomized to prevent the model from over-fitting to specific tool sequence patterns. For evaluation, the data followed the BFCL_V3 framework’s built-in format.

4.2.2

Model Fine-tuning

In this experiment, we selected Qwen2.5-Coder-7B-Instruct as the base model and performed LoRA fine-tuning with four NVIDIA A10 GPUs (24GB VRAM each). Under the default hyper-parameter settings, the batch size was set to 1, with a gradient accumulation step of 16 to simulate large-batch gradient updates under limited computational resources. The learning rate warm-up ratio was set to 0.1, with a peak learning rate of

8.0

×

10

−

5

8.0

superscript

10

5

8.0\times 10^{-5}

8.0 × 10 start_POSTSUPERSCRIPT - 5 end_POSTSUPERSCRIPT

, followed by a cosine learning rate decay to gradually reduce the learning rate after reaching its peak.

The training process used bf16 precision and ran for 10 iterations. For testing, checkpoints were typically selected between the

7

t

⁢

h

superscript

7

𝑡

ℎ

7^{th}

7 start_POSTSUPERSCRIPT italic_t italic_h end_POSTSUPERSCRIPT

and

10

t

⁢

h

superscript

10

𝑡

ℎ

10^{th}

10 start_POSTSUPERSCRIPT italic_t italic_h end_POSTSUPERSCRIPT

iterations. During ablation experiments, the same number of data iterations was maintained to ensure consistency. In the LoRA configuration, the rank size

r

𝑟

r

italic_r

was set to 8, the scaling factor

α

𝛼

\alpha

italic_α

was set to 16, and the dropout rate was set to 0. LoRA fine-tuning was applied to all modules in the model.

During fine-tuning, we used different datasets to train the model, including:

•

DHR_train_1

: Containing 1,260 AI-augmented seed data instances.

•

DHR_train_2

: Containing 1,035 human-augmented seed data instances.

•

DHR_train_3

: A combined dataset containing 1,260 AI-augmented and 1,035 human-augmented seed data instances, totaling 2,295 data instances.

The average final token length of each data entry exceeded 3,000 tokens across these datasets.

The experiment compared the impact of different datasets on the model’s performance after fine-tuning. Furthermore, we examined the performance disparities of models trained on a dataset that combined AI-generated and human-augmented data under different training conditions, and compared their effectiveness with that of generic models.

4.2.3

Model Evaluation

In this experiment, we evaluated the function-calling instructions generated by the model using the AST method. A total of 207 manually annotated seed questions were selected as the

DHR_test_A

test set to compare the effects of different hyper-parameter settings and mixed data on model fine-tuning. Additionally, we designed 135 distinct questions as the

DHR_test_B

test set to prevent evaluation bias caused by potential over-fitting to seed data.

After obtaining the model’s output and its overall accuracy score, we further analyzed the distribution of error types in certain test scenarios. These error types included

Structure Errors (SE)

,

Tool Errors (TE)

,

Parameter Errors (PE)

. These errors are calculated in certain orders during evaluation: Function-calling instructions with structural parsing errors could not be evaluated for tool selection or parameter input accuracy, while instructions with tool selection errors could not be evaluated for parameter input accuracy. The calculation formulas are expressed in the equations below:

P

⁢

(

Structural Completeness Rate

)

=

N

⁢

(

Test Set

)

−

N

⁢

(

SE

)

N

⁢

(

Test Set

)

𝑃

Structural Completeness Rate

𝑁

Test Set

𝑁

SE

𝑁

Test Set

P(\text{Structural Completeness Rate})=\frac{N(\text{Test Set})-N(\text{SE})}{%
N(\text{Test Set})}

italic_P ( Structural Completeness Rate ) = divide start_ARG italic_N ( Test Set ) - italic_N ( SE ) end_ARG start_ARG italic_N ( Test Set ) end_ARG

(4)

P

⁢

(

Tool Selection Acc.

)

=

N

⁢

(

Test Set

)

−

N

⁢

(

SE

)

−

N

⁢

(

TE

)

N

⁢

(

Test Set

)

−

N

⁢

(

SE

)

𝑃

Tool Selection Acc.

𝑁

Test Set

𝑁

SE

𝑁

TE

𝑁

Test Set

𝑁

SE

P(\text{Tool Selection Acc.})=\frac{N(\text{Test Set})-N(\text{SE})-N(\text{TE%
})}{N(\text{Test Set})-N(\text{SE})}

italic_P ( Tool Selection Acc. ) = divide start_ARG italic_N ( Test Set ) - italic_N ( SE ) - italic_N ( TE ) end_ARG start_ARG italic_N ( Test Set ) - italic_N ( SE ) end_ARG

(5)

P

⁢

(

Parameter Filling Acc.

)

=

N

⁢

(

Test Set

)

−

N

⁢

(

SE

)

−

N

⁢

(

TE

)

−

N

⁢

(

PE

)

N

⁢

(

Test Set

)

−

N

⁢

(

SE

)

−

N

⁢

(

TE

)

𝑃

Parameter Filling Acc.

𝑁

Test Set

𝑁

SE

𝑁

TE

𝑁

PE

𝑁

Test Set

𝑁

SE

𝑁

TE

P(\text{Parameter Filling Acc.})=\frac{N(\text{Test Set})-N(\text{SE})-N(\text%
{TE})-N(\text{PE})}{N(\text{Test Set})-N(\text{SE})-N(\text{TE})}

italic_P ( Parameter Filling Acc. ) = divide start_ARG italic_N ( Test Set ) - italic_N ( SE ) - italic_N ( TE ) - italic_N ( PE ) end_ARG start_ARG italic_N ( Test Set ) - italic_N ( SE ) - italic_N ( TE ) end_ARG

(6)

In addition, we conducted confusion matrix evaluations for the fine-tuned model trained on the

DHR_train_3

dataset. We calculated Precision, Recall, and F1 scores for each workflow invoked by the models, providing detailed insights into the model’s performance in workflow selection.

4.3

Experimental Results

Table 3

:

Training Results of Experiments in Digital HR Scenario. The models DHR_train_1_ft, DHR_train_2_ft, and DHR_train_3_ft were trained using their respective datasets.

DHR_test_A

DHR_test_B

Structural

Completeness

Rate (%)

Tool

Selection

Accuracy (%)

Parameter

Filling

Accuracy (%)

Qwen2.5-Coder-

7B-Instruct

22.7

28.1

92.8

82.8

29.6

GPT-4

79.2

88.1

99.0

95.1

89.2

GPT-4o

32.9

34.1

94.2

52.3

66.7

DHR_train_1_ft

85.5

89.6

100

85.5

100

DHR_train_2_ft

94.7

94.1

100

94.7

100

DHR_train_3_ft

97.6

95.6

100

97.6

100

Tab.

3

presents the training results of our experiments. The models DHR_train_1_ft, DHR_train_2_ft, and DHR_train_3_ft were trained using their respective datasets. As shown, DHR_train_1_ft and DHR_train_3_ft significantly outperformed the baseline models Qwen2.5-Coder-7B-Instruct and GPT-4o across both test sets.
To further analyze the experimental results, we categorized the error types for different models on the

DHR_test_A

test set, with the findings summarized in Tab.

3

. The analysis revealed substantial improvements in all aspects after fine-tuning, including the completeness of function-calling structures, the accuracy of tool selection, and the accuracy of parameter filling. After fine-tuning, the models achieved complete accuracy in output structure completeness and parameter filling accuracy.
In terms of tool invocation accuracy, DHR_train_1_ft surpassed both the base model and GPT-4o, DHR_train_2_ft reached the performance level of GPT-4, and DHR_train_3_ft exceeded GPT-4. These findings underscore the effectiveness of fine-tuning for specific scenarios, demonstrating that even small-parameter models can be effectively optimized for specialized tasks, thereby narrowing the performance gap with larger, resource-intensive LLMs.

4.4

Confusion Matrix Evaluation Results

To verify the rationality of these workflow designs and test sets, we also analyzed the performance of the Qwen2.5-Coder-7B-Instruct model after fine-tuning on the

DHR_test_A

test set. The results were represented using confusion matrices in Fig.

3

and relevant performance metrics in Fig.

4

.

(a)

DHR-train-3-ft

(b)

Qwen2.5-Coder-7B-Instruct

(c)

GPT-4

Figure 3

:

Tool Selection Confusion Matrix Heatmap

Figure 4

:

F1 Score Comparison of Different Models across Various Tools. For

[12] Marketing Staff Achievement Inquiry

, the original model failed all four test instances, thus we use zero to represent its F1 score.

The results demonstrate significant improvements in tool selection accuracy after fine-tuning. Compared to the origin model and GPT-4, our fine-tuned model not only achieved near-perfect accuracy in selecting distinctly different tools but also demonstrated improved differentiation between similar tools. Notably, the model showed enhanced F1 score for the three error-prone tools:

[4] Employee Capability Analysis

,

[6] Employee Work Experience

, and

[14] Job Matching

, indicating a comprehensive enhancement of the model’s function selection capabilities within the scenario.

5

Ablation Study

To further investigate the factors affecting the model’s training performance, we conducted a series of in-depth ablation experiments. These experiments involved adjustments to the base model, the composition of training data, data length, and the selection of training hyper-parameters. The tests were conducted on the

DHR_test_A

test set, with the results and analyses detailed below:

5.1

Comparison of Base Models after Training

First, we explored the impact of different base models on training performance. The experiments focused primarily on the Qwen series models, which were trained using the

DHR_train_3

dataset.

Tab.

4

shows that

Qwen2.5-Coder-7B-Instruct

performed the best and is the most suitable base model for this scenario. The Coder model excels in generating function-calling data formats, as its training set includes a higher proportion of code and structured data, which enhances its performance in the function-calling tasks. Qwen1.5-7b-Chat also achieved very good results, likely due to the DPO (Direct Preference Optimization) training it underwent in later stages, making it more adaptable for fine-tuning in downstream tasks

[

rafailov2024direct

]

. In contrast, Qwen2.5-7B, as a pre-trained model without fine-tuning on tasks, did not perform as well as the instruction-tuned models.

Table 4

:

Impact of Different Base Models on Training Performance. Qwen2.5-Coder-7B-Instruct performed the best in Qwen series and is the most suitable base model for this scenario.

Structural

Completeness

Rate (%)

Tool

Selection

Accuracy (%)

Parameter

Filling

Accuracy (%)

DHR_test_A

Qwen2.5-Coder-

7B-Instruct

100

97.60

100

97.60

Qwen2.5-7B-Instruct

100

96.60

100

96.60

Qwen2-7B-Instruct

100

90.80

100

90.80

Qwen1.5-7b-Chat

100

97.10

100

97.10

Qwen2.5-7B

94.20

85.60

44.30

35.70

5.2

Comparison Between AI-Generated and Manually Annotated Seed Data

Table 5

:

Impact of Different Dataset on Training Performance. Each dataset was randomly sampled to include only 1,000 data instances for fine-tuning.

Datasets

DHR_test_A

DHR_test_B

DHR_train_1_1000

83.09

87.41

DHR_train_2_1000

95.17

94.81

DHR_train_3_1000

92.27

96.30

We used Qwen2.5-Coder-7B-Instruct as the base model for subsequent ablation experiments. To compare the impact of different training datasets on model performance, we conducted fine-tuning tasks using the

DHR_train_1

,

DHR_train_2

, and

DHR_train_3

datasets. Each dataset was randomly sampled to include only 1,000 data instances for fine-tuning. The results, shown in Table

5

, indicate that the DHR_train_3_1000_ft model achieved particularly significant improvements in both the

DHR_test_A

and

DHR_test_B

benchmark tests. This demonstrates that training the model using a combination of AI-generated and manually annotated seed data yields better results. One possible explanation is that the AI-generated data expanded the coverage of workflows that were underrepresented in the manually annotated seed data, thereby improving the model’s performance in those areas, and preventing the model from over-fitting to several high-frequency tools. The proportion of corresponding data has been increased, which has played a role in balancing the data set.

5.3

Comparison of Training Data Cutoff Length

Figure 5

:

Comparison of Model Performance with Different Training Data Cutoff Length

In fine-tuning tasks involving unstructured natural language, the fine-tuning framework offers a data cutoff method. By setting a maximum input length for training data, this method processes the data by truncating the parts that exceed the defined length, retaining only the portion within the maximum input limit

[

zheng2024llamafactory

]

. This approach helps prevent excessive consumption of GPU memory. For tasks that use natural language as training data, truncating text is a common practice, as truncating some content typically does not significantly affect model performance. However, since function-calling tasks rely heavily on structured data, truncating training data could have a more substantial impact. To investigate the effect of cutoff length on structured data, we trained models with various cutoff lengths. The results are shown in Figure

5

. The results indicate that as the cutoff length decreases, the training performance deteriorates progressively. Given that the average length of our training data exceeds 3,000 tokens, truncation occurs when the

cutoff_len

is set to 3,000 or less. The results show that the greater the degree of truncation to the origin data, the more significant the negative impact on the model’s final performance. Therefore, when training models for tasks involving structured data, it is crucial to ensure that the truncation length exceeds the average length of the training data to preserve training effectiveness.

5.4

Comparison of Tool Description Length

The length of tool descriptions plays a critical role in data synthesis, model training, and evaluation. Detailed descriptions are highly beneficial for data synthesis and generic model invocation. However, they also lead to longer context lengths, increasing GPU memory consumption and making it more challenging for the model to pay attention to key information. In our experiment, we explored the impact of workflow description length on training performance. Specifically, we categorized description lengths into three types:

•

Long Description:

The original training data containing lengthy and detailed descriptions.

•

Short Descriptions:

Simplified versions of long descriptions, manually reduced to retain only essential functions and key concepts.

•

No Descriptions

: Tool descriptions were completely omitted.

Table 6

:

Examples of Tool Description with Different Length

Description Type

Example

Long Description

{"name": "Marketing_Employee_Data_Inquiry", "description": "Query the

name of marketing employees, year, marketing type; primary department,

total sales revenue, total sales revenue of the year as a percentage of the total

revenue of the primary department, secondary department, total sales revenue

of the year as a percentage of the total revenue of the secondary department,

total sales revenue of the secondary department, year; query the name of the

customer sold by the marketing staff, the customer code, the type of the cust-

omer, the primary industry, the secondary industry, the Chinese brand and so

on; query the marketing staff The name of the virtual agency, including chan-

nel manager, organization and scoring; marketing personnel sales type setting

table can query the sales staff and sales type: composite, channel, customer

and so on."}

Short Description

{"name": "Marketing_Employee_Data_Inquiry", "description": "Inquire

about marketing employee’s name, year, type of marketing, information

about customers sold by the marketing staff, name of the virtual organization,

salesperson and type of sales, and much more."}

No Description

{"name": "Marketing_Employee_Data_Inquiry", "description": ""}

Examples of different description length are shown in Table

6

. Using these variations, we constructed training and test datasets. Models were trained on their respective training sets and tested on all three test sets with different description length. Results are shown in Figure

6

. The

short description model

achieved the best performance on the short description test set, reaching an accuracy of 98.6% on its own test set, and also demonstrated the best generalization across the long and no-description test sets. The long description model showed moderate generalization performance, while the no description model performed reasonably well across all test sets. These results suggest that shortening tool descriptions during training can effectively optimize the model’s overall performance by balancing specificity and generalization.

Figure 6

:

Comparison of Model Performance with Different Tool Description Length

5.5

Comparison of Multi-LoRA Adapter Merging

Table 7

:

Impact of Merging Multiple LoRA Weights.

DHR_train_3

is randomly split in to two training set, based on the seed data groups before augmentation. Similarly, test set is split using the same seed data groups, resulting in the

DHR_test_A_v1

and

DHR_test_A_v2

test sets. The cat (weight=1) merging method achieved the best performance.

DHR_test_A_v1

DHR_test_A_v2

Avg.

LoRA_a

94.2

81

87.6

LoRA_b

92.8

81.9

87.35

LoRA_linear

(weight=0.5)

91.3

78.3

84.8

LoRA_linear

(weight=1.0)

96.7

83.7

90.2

LoRA_dare_linear

(density=0.8)

94.7

81

87.85

LoRA_cat

(weight=0.5)

94.7

82.8

88.75

LoRA_cat

(weight=1.0)

96.6

84.2

90.4

LoRA_cat

(weight=2.0)

0

0

0

LoRA_svd

94.7

82.8

88.75

LoRA_ties

(density=0.8)

91.3

78.3

84.8

LoRA_ties_svd

(density=0.8)

94.7

82.8

88.75

To explore a training-free method, we studied the impact of merging multiple LoRA weights, aiming to enable the model to continuously incorporate new data while mitigating the effects of outdated data. Results are shown in Table

7

. Compared to continual learning, LoRA merging provides greater flexibility as a plug-in mechanism, allowing the model to adapt to different scenarios by selecting data from appropriate cycles. For this experiment, the

DHR_train_3

dataset was randomly split into two parts (split by seed groups) to simulate new data generated at different time periods. These were used to train two separate LoRA models (

LoRA_a

and

LoRA_b

) using identical training methods. After training, we attempted various merging methods to combine the two LoRA weights and tested their effectiveness on the

DHR_test_A

test set. The merging methods included

linear

,

cat

,

dare_linear

,

svd

,

ties

, and

ties_svd

[

peft

]

. The results, shown in the Table

7

, indicate that the cat(

weight=1

) merging method achieved the best performance. The merged LoRA outperformed both

LoRA_a

and

LoRA_b

, demonstrating superior generalization capabilities. This highlights the potential of LoRA merging not only to improve model adaptability but also to provide a flexible and efficient approach for managing data life-cycle effects in enterprise scenarios.

6

Conclusion

In this paper, we introduced a specialized training pipeline for function-calling capabilities in professional scenarios. This pipeline encompasses the generation and augmentation of scenario-specific function-calling data, LoRA fine-tuning of models, and comprehensive evaluation and analysis of function-calling performance. The experimental results demonstrate that our pipeline can improving the accuracy of generic function-calling models in specialized scenarios. Notably, a 7B-parameter model trained using this pipeline achieved accuracy levels superior to state-of-the-art LLMs. This is particularly significant for enterprises, as the pipeline enables the automated training of intelligent agents’ core engines tailored to various business scenarios.

Beyond the digital HR scenario explored in this study, the pipeline can be applied to other contexts, such as integrating with edge models in edge devices, supporting diagnostic systems in healthcare, or powering super employees within organizations. Furthermore, with the inclusion of a data feedback module, the system could utilize data generated from users’ interactions with agent applications. By leveraging AI and the designed pipeline, the returning data can be automatically annotated and used for continual model training, enabling iterative updates.

7

Limitation & Future Work

Currently, the proposed pipeline has certain limitations. First, it focuses solely on single-turn data synthesis and does not account for more complex scenarios, such as multi-step and multi-turn interactions, which are highly practical in real-world applications. Additionally, the current pipeline’s data filtering process for generated data is relatively simplistic. Future efforts will aim to enhance the filtering of augmented data to further improve data quality. Despite these limitations, we believe this pipeline represents a step forward in applying LLMs to real-world enterprise scenarios.

\printbibliography

Appendix Appendix A

Data Synthesis Prompt Templates

Prompt Templates are translated to English.

Appendix A.1

Question Generation Prompt Template

⬇

prompt

=

f

""

"You

will

act

as

a

user

asking

questions,

using

the

provided

function

parameter

list

{key}:{params_data}

to

generate

questions.

Here,

{key}

represents

the

function

name,

and

the

function

description

is

{function_description}.

Follow

the

requirements

below:

-

Generate

{number}

questions

for

each

function.

-

Include

names,

ID

strings,

and

ID

parameters

in

each

question,

ensuring

that

the

ID

names

are

retained.

-

When

the

question

involves

employee

IDs,

you

may

randomly

generate

human

names.

-

The

ID

is

known

information

and

must

be

shown

in

the

question.

-

Do

not

include

the

function

name

in

the

question.

-

Only

generate

the

questions,

without

answers

or

any

additional

content.

Output

format

example:

1.

Question

content

2.

Question

content

...

5.

Question

content"

""

Appendix A.2

Question Generation with real name

⬇

prompt

=

f

""

"You

are

a

data

annotator.

Your

task

is

to

generate

a

set

of

diverse

questions

for

the

given

function.

The

constructed

questions

will

be

directly

used

as

parameters

for

function

calls,

demonstrating

how

these

functions

can

be

applied

in

real-world

scenarios.

Please

adhere

to

the

following

requirements:

-

Ensure

the

questions

are

as

diverse

as

possible

and

not

limited

to

the

"

example

"

format.

-

Project

names,

course

names,

etc.,

can

be

reasonably

fabricated.

Ensure

the

following:

-

Provide

a

variety

of

questioning

styles

and

tones,

consistent

with

those

used

by

various

types

of

consulting

professionals.

-

The

first

half

of

the

questions

should

consist

of

**{half}

single

queries**

(e.g.,

"

xxxxxx

?

"),

and

the

second

half

should

consist

of

**{half}

multi-part

queries**

(e.g.,

"

xxxxxx

?

xxxxxx

?

xxxxxx

?

").

-

Generate

questions

based

on

the

function

tool’s

description

and

usage

but

avoid

explicitly

mentioning

the

function

name.

-

Include

characters

(IDs

or

names)

in

the

questions,

ensuring

that

all

IDs

or

names

are

selected

from

the

following

list:

**{ids}**.

-

Generate

each

question

on

a

single

line.

Do

not

include

any

bullet

points,

numbering,

or

blank

lines.

The

response

should

contain

only

the

generated

questions,

without

any

additional

content.

Based

on

the

above

instructions

and

examples,

generate

**{number}

different

questions**

for

the

function

**"

{

func_name

}

"**.

The

detailed

function

description

is

as

follows:

**{func_desc}**

Now,

generate

a

total

of

**{number}

different

questions**

following

the

format

above.

Ensure

the

number

of

questions

is

accurate.

"

""

Appendix A.3

Question Generation without name

⬇

prompt

=

f

""

"You

are

a

data

annotator.

Your

task

is

to

generate

a

set

of

diverse

questions

for

the

given

function.

The

constructed

questions

will

be

directly

used

as

parameters

for

function

calls,

demonstrating

how

these

functions

can

be

applied

in

real-world

scenarios.

Please

adhere

to

the

following

requirements:

-

Ensure

the

questions

are

as

diverse

as

possible

and

not

limited

to

the

"

example

"

format.

-

Project

names,

course

names,

etc.,

can

be

reasonably

fabricated.

Ensure

the

following:

-

Provide

a

variety

of

questioning

styles

and

tones,

consistent

with

those

used

by

various

types

of

consulting

professionals.

-

The

first

half

of

the

questions

should

consist

of

**{half}

single

queries**

(e.g.,

"

xxxxxx

?

"),

and

the

second

half

should

consist

of

**{half}

multi-part

queries**

(e.g.,

"

xxxxxx

?

xxxxxx

?

xxxxxx

?

").

-

Generate

questions

based

on

the

function

tool’s

description

and

usage

but

avoid

explicitly

mentioning

the

function

name.

-

Avoid

including

specific

individuals

in

the

questions;

instead,

focus

on

companies

or

groups

of

people.

-

Generate

each

question

on

a

single

line.

Do

not

include

any

bullet

points,

numbering,

or

blank

lines.

The

response

should

contain

only

the

generated

questions,

without

any

additional

content.

Based

on

the

above

instructions

and

examples,

generate

**{number}

different

questions**

for

the

function

**"

{

func_name

}

"**.

The

detailed

function

description

is

as

follows:

**{func_desc}**

Now,

generate

a

total

of

**{number}

different

questions**

following

the

format

above.

Ensure

the

number

of

questions

is

accurate.

"

""

Appendix A.4

Short Question Generation

⬇

prompt

=

f

""

"You

are

a

data

annotator.

Your

task

is

to

perform

data

augmentation

on

the

given

question

based

on

the

function

description,

expanding

the

phrasing

of

the

question.

Question

to

be

expanded:

"

{

question

}

"

Please

follow

these

guidelines:

-

Use

various

modification

methods,

including

but

not

limited

to

changing

the

order

of

elements,

question

style,

tone,

and

omissions.

-

You

may

expand

or

simplify

the

question.

-

Subject,

verb,

and

object

can

be

rearranged,

such

as

placing

names

at

the

end

of

the

sentence.

-

If

the

question

includes

characters,

you

may

replace

the

names

in

the

enhanced

versions

with

names

or

IDs

from

the

following

list:

{ids}.

Ensure

that

the

original

meaning

of

the

question

remains

unchanged.

Ensure

that

the

enhanced

versions

are

still

short

questions.

Ensure

that

only

the

final

{number}

answers

are

output,

with

each

question

on

a

separate

line.

Do

not

include

additional

content

or

numbering.

Based

on

the

examples

and

instructions

above,

expand

the

question

into

a

total

of

{number}

different

versions.

Ensure

the

correct

number

of

outputs.

"

""

Appendix Appendix B

Data Augmentation Prompt Templates

Appendix B.1

Question Replacement

⬇

prompt

=

f

""

"

You

are

now

a

data

annotation

expert

capable

of

rewriting

questions

according

to

specific

requirements.

Note:

Directly

respond

with

the

rewritten

questions.

Generate

only

the

rewritten

answers

without

any

extra

content

or

numbered

bullet

points.

Please

follow

these

instructions:

-

Carefully

review

the

contents

in

**{info_dict}**,

and

ensure

that

the

rewritten

questions

align

with

the

descriptions

provided.

-

Replace

key

information

and

content,

such

as

names,

dates,

various

labels,

locations,

etc.

-

Replace

any

names

appearing

in

the

question,

e.g.,

Wang

XX,

with

other

names

from

the

following

list:

**{names}**.

Ensure

all

names

are

replaced.

-

Replace

any

departments

appearing

in

the

question,

e.g.,

Marketing

Department,

with

other

departments

from

the

following

list:

**{departments}**.

-

Replace

any

cities

appearing

in

the

question,

e.g.,

Beijing,

with

other

cities

from

the

following

list:

**{cities}**.

-

Replace

years

mentioned

in

the

question

with

one

of

the

following:

**2021,

2022,

2023**.

-

Replace

any

other

numbers

mentioned,

e.g.,

Issue

344,

with

random

numbers.

-

Replace

any

zodiac

signs,

e.g.,

Sagittarius,

with

others

such

as

Capricorn,

Leo,

Taurus,

Virgo,

Aquarius,

etc.

-

Replace

any

project

amounts

mentioned

with

random

values,

e.g.,

100K,

500K,

2M,

etc.

-

Replace

any

job

positions,

e.g.,

Product

Manager,

with

others

such

as

Data

Analyst,

Marketing

Specialist,

Finance

Manager,

HR

Specialist,

General

Manager,

etc.

-

Replace

any

skills

mentioned,

e.g.,

Big

Data

Development,

with

others

such

as

Software

Development,

Foreign

Language

Proficiency,

Project

Management,

Organizational

Coordination,

Planning

Skills,

etc.

-

Replace

any

technologies

mentioned,

e.g.,

Deep

Learning,

with

others

such

as

Data

Governance,

Large

Model

Fine-Tuning,

Prompt

Engineering,

Machine

Learning,

Computer

Vision,

etc.

-

Ensure

the

rewritten

questions

remain

simple

and

focus

on

core

information.

Avoid

using

prepositions,

conjunctions,

modal

particles,

or

honorifics.

Do

not

include

quotation

marks.

Rewrite

the

question

**"

{

question

}

"**

in

**{number}

different

ways**.

"

""

Appendix B.2

Question Rewriting

⬇

prompt

=

f

""

"You

are

now

a

linguist

capable

of

rewriting

user

queries.

Your

task

is

to

change

the

grammatical

structure

and

questioning

style

of

the

sentence

while

preserving

its

original

meaning

and

key

information.

Key

information

is

referenced

from

{info_dict}.

Please

rewrite

the

given

question

into

{number}

different

forms,

altering

the

syntax

and

questioning

style

without

adding

new

information

or

omitting

any

original

content.

-

Each

rewritten

question

should

be

on

a

separate

line.

-

Do

not

include

any

additional

content,

quotation

marks,

or

numbering.

-

The

rewritten

questions

will

be

directly

used

as

parameters

for

function

calls,

showcasing

how

these

functions

are

applied

in

real-world

scenarios.

Rewrite

the

following

question:

{question}

"

""

Appendix B.3

Question Simplification

⬇

prompt

=

f

""

"You

are

a

data

annotator.

Your

task

is

to

simplify

the

given

question

based

on

the

tool

function

description.

The

question

you

need

to

simplify

is:

"

{

question

}

".

Please

adhere

to

the

following

requirements:

-

Simplify

the

question

into

a

very

concise

phrase.

-

Do

not

change

the

original

meaning

of

the

question.

Example:

Original

question:

"

What

is

Zhang

San

’s

current

salary?"

Simplified:

"Zhang

San

salary"

Provide

one

question

per

line.

Do

not

include

any

additional

content,

quotation

marks,

or

bullet

points.

Based

on

these

examples

and

the

instructions

above,

generate

a

total

of

{number}

different

simplified

questions.

Ensure

the

number

of

questions

is

correct.

"""’

Appendix B.4

Question Error Introduction

⬇

prompt

=

f

""

"You

are

now

a

data

annotator

tasked

with

simulating

user

input

errors

by

modifying

the

original

query.

Make

the

following

types

of

changes:

-

rewrite

the

query

"

{

question

}

"

to

produce

**{number}

rewritten

versions**.

-

Replace

one

character

in

the

user

query

with

a

phonetically

similar

typo.

-

Replace

some

words

with

incomplete

or

scrambled

versions

of

the

words.

-

Repeat

some

characters

within

the

query.

-

Respond

directly

with

the

rewritten

queries

and

do

not

include

any

extra

content.

-

Provide

each

rewritten

query

on

a

separate

line.

-

Do

not

include

quotation

marks,

numbering,

or

any

additional

information

in

your

response.

-

Ensure

the

number

of

rewritten

queries

matches

the

requested

count.

Example:

Original

query:

"

How

many

technical

staff

members

are

in

the

company

?

"

Rewritten:

"

How

many

technical

stuff

members

are

there

in

most

companie

?

"

Follow

these

examples

and

instructions

to

rewrite

the

query.

Ensure

the

quantity

of

responses

is

accurate.

"

""