Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning

Ming Li, Yong Zhang, Shwai He, Zhitao Li, Hongyu Zhao, Jianzong Wang, Ning Cheng, Tianyi Zhou

Introduction

The landscape of natural language processing has witnessed a transformative change with the introduction of large language models (LLMs) like GPT-3 Brown et al. (2020), GPT-4OpenAI (2023), LLaMA Touvron et al. (2023a, b), Mistral Jiang et al. (2023), etc Penedo et al. (2023); Scao et al. (2022). These models offer unprecedented capabilities in generating contextually rich, coherent, and often creative text. A key advantage of them is the ability to follow instructions, which leads to promising performance on zero-shot (prompting) or few-shot (in-context learning) tasks. To achieve it, a supervised learning technique named instruction tuning Wei et al. (2022); Longpre et al. (2023a) is commonly utilized, which finetunes an LLM to produce preferred responses on a broad range of tasks described by natural language instructions.

Earlier works of instruction tuning focus on creating large, varied, and high-quality datasets of various tasks with responses curated by human experts Khashabi et al. (2020); Ye et al. (2021); Wei et al. (2022); Wang et al. (2022); Du et al. (2022), which can be bottlenecked by the intensive human labor. An alternative is to generate the data by a powerful teacher LLM Wang et al. (2023b); Taori et al. (2023); Xu et al. (2023); Li et al. (2023a) but the quality is hard to control and largely depends on the teacher. LIMA Zhou et al. (2023) finds that a mere 1,000 human-crafted high-quality data could significantly improve an LLM’s instruction-following capability, based on which they posit that LLMs acquire most knowledge during the pretraining and thus a few data suffices for instruction tuning.

To further free the human labor in data curation and accelerate the instruction tuning process, a line of recent works apply an extra filter algorithm to select data from the existing dataset. However, the model used in the filtering process usually needs to be as powerful as ChatGPT Chen et al. (2023b); Lu et al. (2023) or requires additional reward data training Du et al. (2023); Bukharin and Zhao (2023), or is the student model itself Li et al. (2023b), which leads to additional expensive cost and latency due to the inference on these large filter models, especially when the original dataset is large while only a tiny fraction of data needs to be selected. These paradigms are presented in Figure 1 (a) and (b). To reduce the filtering cost, we study Superfiltering: Can we use a smaller and weaker model as a filter to select instruction-tuning data for training a larger and stronger model? This was first studied for training small classification models by Coleman et al. (2020) while the effectiveness on the open-domain instruction dataset is un-explored. Recently, Weak-to-Strong Generalization Burns et al. (2023) proposes to utilize a weaker ChatGPT to generate data used to finetune a stronger GPT4 model, which shares a similar spirit with our Superfiltering as depicted in Figure 1(c).

In Superfitering, we find that a smaller and weaker GPT-2 (124M) Radford et al. (2019) suffices to replace previously used large filter models and select high-quality instruction tuning data used to finetune a much larger LLaMA2 (7B or 13B). This is motivated by our main discovery of filter models’ consistency on two data statistical metrics, perplexity and instruction-following difficulty (IFD) score Li et al. (2023b). Despite the differences in scales across different filter models, their rankings of the same instruction tuning dataset are surprisingly consistent, as demonstrated by the large rank correlation coefficients evaluated on different models and datasets. Our thorough empirical study implies that weaker language models possess a capability consistent with their stronger counterparts in comprehending and discerning the difficulty of diverse instructions, though they may differ in other skills like reasoning and generalization.

In extensive experiments, our Superfiltering strategy using GPT-2 as the filter, as exemplified on several widely used instruction datasets, brings significant speedups to data filtering for instruction tuning. By utilizing only 5% of the original data volume, Superfilter allows us to attain LLMs comparable, and in some instances superior, to those achieved by training with full data. Our main contributions can be summarized in three folds:

Weak-to-Strong Consistency on Data Filtering: We discover a strong consistency between small and large LLMs in perceiving and evaluating the difficulty of instruction tuning data.

Efficent Superfiltering Strategy: We propose the first method of Superfiltering that utilizes a small LM, e.g., GPT-2 (124M), to select data for instruction tuning, and brings significant speedups to LLM finetuning pipeline.

Efficacy of Selected Training Data: Superfiltering is precise in allocating high-quality and informative data improving LLM instruction tuning.

Problem Formulation

We define a dataset as DD, containing nn triplets x=(Instruction,[Input],Response)x=(Instruction,[Input],Response) as the instruction tuning data samples. Earlier instructing tuning samples mostly contain separated instructioninstruction and inputinput segments of better controls Wang et al. (2022); Longpre et al. (2023b); Taori et al. (2023), while most of the current datasets directly merge the inputs to instructions Zhou et al. (2023); Chiang et al. (2023); Xu et al. (2023); Li et al. (2023a). For simplicity, we define x=map(Instruction,[Input])x=map(Instruction,[Input]) as the complete instruction and yy as the corresponding response. The mapping function could be the simple concatenation with some control tokens. Thus D={(x1,y1),(x2,y2),…,(xn,yn)}D=\{(x_{1},y_{1}),(x_{2},y_{2}),\ldots,(x_{n},y_{n})\} represents a collection of nn instruction-response pairs.

In the instruction tuning setting, the model is trained to maximize the likelihood of response given the corresponding instruction as the condition. Hence, perplexity can be a potential metric to measure the difficulty. Specifically, the perplexity of a given sample (xi,yi)(x_{i},y_{i}) is defined as:

where NN is the length of response yiy_{i} and yi,jy_{i,j} represents the jjth token in the response yiy_{i}.

Li et al. (2023b) firstly proposes a self-guided method in which no extra models are utilized but needs to calculate Instruction-Following Difficulty (IFD) scores based on the pre-experienced LLM or original pre-trained LLM. The IFD score is a pure statistical metric, that compares the losses or perplexities when the model generates a response yiy_{i} with and without instructional context x1x_{1}, measuring how much help the instruction provides to the generation of the corresponding response. A higher IFD score, indicating less instructional help, suggests a greater difficulty. On the contrary, the low IFD score represents that the given instruction can directly benefit the language model largely even without further training, representing the easiness and necessity of the instruction. For a given instruction-following data pair, the IFD score is calculated as follows:

where PPL(yi∣xi)\text{PPL}(y_{i}|x_{i}) and PPL(yi)\text{PPL}(y_{i}) denote the perplexities of the given model in fitting response yiy_{i} with and without the instruction xix_{i}, respectively.

2 Formulation and Motivations

Superfiltering aims to find a data filtering score (1) that excels in identifying high-quality and informative training data, and (2) computed by a small and low-cost filter model without further training. To this end, we try to find a data evaluation metric consistent between weak and strong language models.

Given a candidate score, we investigate whether it is possible to utilize a much weaker language model, e.g. GPT-2, to calculate for the relatively stronger student model. We hypothesize that, although the intrinsic abilities between weak and strong language models vary dramatically, indicated by the discrepancies of perplexities on the pretraining stage, their ability to perceive instruction difficulty could be similar. To verify our hypothesis, experiments are conducted and presented in Section 3.

To verify the hypothesis, we conduct a thorough empirical study of the consistency of perplexities computed by different language models on the same instruction-tuning dataset. In Section 3.1, we focus on verifying the consistency of perplexity across weak-to-strong models by comparing their scale and orderings of samples on each dataset. The results show that though the scales vary drastically, the orderings remain consistent, which verifies our hypothesis. In Section 3.2, we conduct the same study on IFD scores, on which both the scales and the orderings are consistent across weak-to-strong models, indicating IFD score as a more promising score for Superfiltering than perplexity.

Figure 2 compares each Superfiltering-selected-data finetuned model and the full-data finetuned model by using GPT-4 as a judge to decide their numbers of wins/ties/losses on a test set of instructions. More details of the evaluation metric can be found in Section 4.3. Superfiltering-trained models always outperform the baseline given different base models, datasets, and selection ratios, demonstrating the effectiveness of our proposed weak-to-strong Superfiltering scheme.

Weak-to-Strong Consistency

In this section, we delve into the hypothesis that weak and strong language models share a relatively consistent capability in perceiving the difficulties of instruction tuning samples.

As mentioned in the previous section, we first need to have a grasp of to what extent language models of different sizes are consistent with each other in understanding instructions and generating corresponding responses. Thus we calculate the perplexity scores of several pretrained language models, including relatively small language models like GPT-2 (124M), GPT-2-large (774M), GPT-2-XL (1.5B) Radford et al. (2019), GPT-NEO (1.3B) Black et al. (2021) and recent relatively strong LLaMA2-7b Touvron et al. (2023b), on several instruction-tuning dataset including the Alpaca Taori et al. (2023), Alpaca-GPT4 Peng et al. (2023), and WizardLM 70k Xu et al. (2023). The results are shown in Figure 3 (upper), where each box presents a perplexity distribution of a given dataset and language model. A clear tendency can be found that the stronger the language models are, the lower the perplexities are, which is consistent with the common beliefs for LLM pretraining: the better a language model, the lower this perplexity.

The above experimental results only showcase the perplexity scales of different models and neglect the potential perplexity ordering/ranking of different data samples, which is much more vital for data filtering. Thus to evaluate the similarity in perplexity ordering on a given dataset between different models, Spearman’s rank correlation coefficient (Spearman’s ρ\rho) is utilized. Spearman’s ρ\rho is a non-parametric measure used to assess the strength and direction of the relationship between two variables that are ranked or ordinal in nature. For two lists containing the same elements but different ordering, this value measures the similarity of the ordering in the range of −1-1 to 11. The closer the value is to 11, the more consistent the ordering of these two lists.

Specifically, within each dataset DD, we sort the data samples based on the perplexity scores calculated by different models, resulting in several lists containing the same data but different orders, noted as DPPL, GPT-2D_{\text{PPL, GPT-2}}, DPPL, LLaMA2-7BD_{\text{PPL, LLaMA2-7B}}, etc. Since most of the fine-tuned experiments in our work are implemented on the LLaMA2-7B model, we set DPPL, LLaMA2-7BD_{\text{PPL, LLaMA2-7B}} as our standard sorted list and calculate the Spearman’s ρ\rho between the sorted lists of small language models and LLaMA2-7B:

where gg is the function of calculating this coefficient. All the resulting values on different instruction tuning datasets and different models are presented in Table 1, the Spearman’s Coefficient-Perplexity column.

From the results, we can see even the lowest coefficient value is still greater than 0.70.7, calculated between GPT-2(124M) and LLaMA2-7B, and the highest coefficient value is greater than 0.850.85, calculated between GPT-NEO(1.3B) and LLaMA2-7B. The values presented in the table are reasonably high, indicating the consistent capability of different models in perceiving instructions. Moreover, there is also a clear tendency that the stronger the language models are, the higher the coefficient values. Comparing the perplexity distributions in Figure 3 (upper) and the coefficient values in Table 1, a clear consistency can be revealed: Despite the large variance in the scales of perplexities generated by different language models, representing the intrinsic abilities of different language models, the high consistency in the perplexity ordering indicates the similarity of them to understand instructions. That is to say, for a given instruction tuning sample, if the weak language models find it hard to generate based on the corresponding instruction, the strong models might probably feel the same way even though their probability of generating this response is much larger, and vice versa.

This phenomenon directly provides a glance at the weak-to-strong perplexity consistency, which serves as the basis for utilizing weak language models as the proxies for strong language models.

2 Weak-to-Strong IFD Consistency

Though a clear consistency in the perplexities of different language models is revealed by the above experiments, the perplexity does not directly represent the difficulty or quality of the instruction tuning sample and is thus not able to be used for the data section. Thus we further extend our findings to the Instruction-Following Difficulty (IFD) score proposed by Cherry LLM Li et al. (2023b). It is used to select a subset of high-quality samples from the given instruction-tuning dataset to train an LLM with better performance.

Similarly, we calculate the IFD scores on different instruction-tuning datasets with different language models and draw their distributions as shown in Figure 3 (lower). We observe that though the perplexity scales vary noticeably between models, the IFD scales remain similar, indicating its potential to be the general selection metric for different models. Furthermore, the IFD-based Spearman’s ρ\rho are also presented in Table 1 Spearman’s Coefficient-IFD score column. Similar to the perplexity-based coefficient values, IFD-based values also remain high, indicating a strong consistency of IFD rankings calculated on different models. Such a consistency validates the scalability of weaker models in evaluating instruction difficulty, indicating their adeptness at identifying complex instructions akin to their stronger counterparts. Another interesting phenomenon is that the IFD-based coefficient values are greater than perplexity-based values on high-quality datasets, e.g. Alpaca-GPT4 and WizardLM 70k, indicating an even higher consistency in IFD scores for these datasets.

To provide an even further apparent glance at this consistency, we calculate the overlap ratio when utilizing IFD scores to select the high-quality subset. The performances of the LLMs could be slightly estimated by the overlap ratio due to the previous success of this metric. As the percentage threshold increases from 5%5\% to 15%15\%, there is a significant and growing overlap in the samples identified by the weaker models and strong models like the LLaMA2-7B model. Although the overlap is not complete, it is substantial, this increasing overlap with higher thresholds reinforces our hypothesis, affirming a consistent and scalable capability in instruction evaluation across models of varying sizes.

This weak-to-strong IFD consistency directly verifies our hypothesis that language models with different sizes possess similar capabilities in understanding the difficulty of the instructions, even though their intrinsic abilities are varied. It means that the difficult instruction tuning samples defined by the IFD scores are probably “generally” difficult no matter what language model is utilized for the calculation. This phenomenon directly makes it possible to utilize weak language models as the proxies for strong language models for calculating the IFD scores, and thus, to select data for instruction tuning.

3 Superfiltering

From the above section, we observe that the IFD score is a highly consistent metric when calculating based on different instruction-tuning datasets and varied-size language models. Thus we propose “Superfiltering”, the first approach utilizing only small language models, i.e. GPT-2 (124M) Radford et al. (2019) to filter data for the instruction tuning of modern LLMs. Superfiltering uses smaller, less resource-intensive models (referred to as “weak” models) as effective substitutes for larger models (referred to as “strong” models) in the data evaluations. For the first time, making this process so efficient as to put it into practical usage. Specifically, following Li et al. (2023b), for the given instruction-tuning dataset, the GPT-2 model is directly used to calculate the IFD score of each sample. Then the top kk-percent samples with the highest IFD scores under 11 are selected for faster instruction tuning.

Experimental Setup

The Alpaca dataset Taori et al. (2023) is developed by Stanford University, comprises 52,000 instruction-following samples, and was created using the self-instruct paradigm Wang et al. (2023b). This dataset was generated by leveraging OpenAI’s text-davinci-003 model. The Alpaca dataset represents a classical dataset with moderate qualities, to further verify our method on the originally high-quality dataset, we also implement our method on the Alpaca-GPT4 dataset Peng et al. (2023), which contains the responses generated by GPT4.

2 Implementation Details

We utilize the prompt and code base from Vicuna Chiang et al. (2023) and flash attention Dao et al. (2022) while the overall training arguments are aligned with the common training configuration. The Adam optimizer Kingma and Ba (2017), with a 2×10−52\times 10^{-5} learning rate for the LLaMA2-7B model Touvron et al. (2023b) and a 1×10−51\times 10^{-5} learning rate for the LLaMA2-13B model, and a batch size of 128128, steer the training across three epochs with a max length of 20482048. The warmup rate is set to 0.030.03.

3 Evaluation Metrics and Benchmarks

Evaluating responses generated by Large Language Models (LLMs) like GPT-4 remains a complex and ongoing research area, particularly for open-domain questions where establishing a clear ground truth is challenging. Traditional methods often fall short in assessing the instruction-following ability of these models. Recent trends, however, involve using LLMs themselves, such as GPT-4, as evaluators, a practice that has gained widespread acceptance in the field Touvron et al. (2023b); Chiang et al. (2023); Dettmers et al. (2023); Liu et al. (2023). Previous studies Zheng et al. (2023); Li et al. (2023c); Sottana et al. (2023) have shown that GPT4’s evaluations are consistent with human evaluations. We utilized the testing instruction set from WizardLM Xu et al. (2023) and Vicuna Chiang et al. (2023) which contain 218218 and 8080 diverse human-curated instructions respectively.

Our study adopts the evaluation strategy as outlined by Chen et al. (2023b); Li et al. (2023b, a), involving a detailed rating system for model-generated responses. Each response is scored reflecting various dimensions such as the accuracy and relevance of the response. This method is in line with previous research efforts to assess the effectiveness of language models more accurately. Moreover, to address the issue of positional bias, as discussed in the works of Ko et al. (2020); Wang et al. (2023a), we present the responses generated by the model in two separate sequences for evaluation by the LLM judge. This approach aims to ensure a more balanced and unbiased assessment of the model’s performance. Then for each instruction, we compare the responses by "Win-Tie-Loss".

3.2 AlapcaEval Leaderboard

The AlpacaEval Leaderboard, utilizing the AlpacaFarm Dubois et al. (2023); Li et al. (2023c) evaluation dataset, is an automated, efficient, and reliable evaluation tool for LLMs. It benchmarks LLMs’ performance in following generic user instructions by comparing their outputs with those from Davinci003, demonstrating high alignment with human expert annotations. AlpacaFarm, underlying AlpacaEval, is a cost-effective simulator for research on learning from human feedback, significantly reducing the time and cost traditionally associated with such studies. While AlpacaEval offers valuable insights, it primarily focuses on simpler instructions and does not encompass safety evaluations or complex tasks, and its evaluation may correlate win rates with response lengths. These tools represent significant advancements in LLM evaluation and development, enabling more accessible and diverse research. Considering our budget, we only run the evaluation on 5%5\% settings.

3.3 Open LLM Leaderboard

The Huggingface Open LLM Leaderboard, incorporating the evaluation method from the Eval Harness Gao et al. (2021), serves as a comprehensive framework for evaluating generative language model capabilities. It focuses on four critical benchmarks: ARC Clark et al. (2018), HellaSwag Zellers et al. (2019), MMLU Hendrycks et al. (2021), and TruthfulQA Lin et al. (2022). These benchmarks test the models on various aspects, such as reasoning, common-sense understanding, and factual accuracy. The leaderboard offers an effective platform for comparing different LLMs, providing valuable insights into their performance across these diverse and challenging tasks

Experimental Result

In this section, we present the evaluation results of three different evaluation settings as described in the previous section as shown in Table 2. The Pair-Wise Winning Score indicates the result directly comparing with the corresponding model trained with full data. These values that are greater than 1.01.0 represent better responses generated by our Superfiltering models than full data models. The detailed win-tie-lose numbers are presented in Figure 2. Moreover, the performance of our models and baseline models on the Huggingface Open LLM Leaderboard and the AlpacaEval Leaderboard are also presented in Table 2 where we can see our models using 5%5\%, 10%10\%, 15%15\% data outperform the models trained with full data on both benchmarks on both LLaMA2-7B and LLaMA-13B settings. These results further showcase the effectiveness of our automatically selected data. Moreover, the usefulness of Superfiltering on the high-quality Alpaca-GPT4 dataset further shows the potential of our method, which is astonishing that a pure statistical metric based on a weak language model like GPT-2 is able to filter the responses generated by GPT-4.

2 Ablation Study

In this subsection, extensive ablation experiments are conducted to validate the effectiveness of our Superfiltering. The experiments are performed on the LLaMA2-7B model using the Alpaca dataset. Our focus is on two aspects: the impact of different data selection strategies and the effect of using various language models for data selection. All models are trained under the same settings.

As shown in Table 3, in addition to our method “Superfiltering (GPT-2)”, we also try several baseline strategies: “Random” represents the models trained with randomly selected data. “Diversity” represents the models trained with data considering only diversity, by utilizing the k-means algorithm. “Perplexity” represents the models trained with data based on the perplexity calculated on GPT-2. Moreover, the lower part of the table lists the models using the IFD score to select the training subset, powered by other language models. The performances of models are assessed by the pair-wise winning score, which is calculated as (Num(Win)−-Num(Lose))//Num(All) +1+1, and all the comparisons are performed by GPT4 on the WizardLM test set.

As shown in Table 3, compared with other strategies, models trained with our method consistently outperform the models trained on the full dataset, indicating the efficacy of our method. Regarding the impact of different language models, whichever language model is utilized to calculate the IFD scores, the corresponding models would surpass the baseline model, indicating the strong consistency and transferability of the IFD score as the selection metric. Moreover, the models using LLaMA2-7B reasonably achieve the highest performance, due to the consistency between the model to calculate the IFD scores and the model to be trained.

Further Discussion

In the realm of data selection for language model instruction tuning, our Superfiltering introduces a transformative advantage: the unnecessity of training for even weak language models. Traditional proxy-based methods like Coleman et al. (2020) and Nguyen et al. (2022) are required to further train weak models to bridge the performance gap with stronger models. However, our study reveals that pre-trained weak models are naturally effectively capable of acting as the proxies for strong models when utilizing the IFD for data selection, without requiring any additional training.

Moreover, in the context of instruction tuning data selection, model training is always necessary if no extra strong models like ChatGPT or other trained reward models are utilized. Lu et al. (2023) utilizes chatGPT to tag the instruction datasets and train LLMs for tagging instruction samples based on these data. Wei et al. (2023) first splits a range of subsets from the original data and records each fine-tuned model’s performance on the validation set as the labels of data quality. Then a self-attention network as the data selector is trained. Cao et al. (2023) utilizes a trained regression model to estimate the inference losses as data qualities on several datasets.

The performances of the resulting efficient selection models are appealing while a possible concern is the generalizability since they all need extra performance indicators, i.e. development set, in existing datasets. On the contrary, our approach directly utilizes established, widely accepted models like GPT-2, which are known for their broad applicability and generalization ability. These models do not necessitate fine-tuning on specific datasets, thereby reducing the risk of out-of-distribution issues. Our method demonstrates that these smaller models can be deployed in a “plug-and-play” manner, achieving commendable performance immediately. This innovative approach not only simplifies the data selection process but also revolutionizes the efficiency and applicability of such methods in large language model instruction tuning.

Related Work

Recent advancements in natural language processing (NLP) have been significantly influenced by instruction tuning, a method that tailors large language models (LLMs) for varied tasks using explicit instructions Wei et al. (2022); Sanh et al. (2022); Longpre et al. (2023a). This approach has enhanced LLMs’ ability to understand and follow instructions in diverse contexts. Initial research in this area focused on expanding dataset sizes to improve instruction-following capabilities Honovich et al. (2023); Wang et al. (2023b). Additionally, innovative methods, such as using LLMs to generate instructional data, are being explored to streamline and enhance the instruction tuning process Wang et al. (2023b); Xu et al. (2023); He et al. (2023); Li et al. (2023a).

2 Instruction Tuning Data Selection

To further select the data for more efficient instruction tuning, existing automatic data selection methods mainly utilize extra LLMs for the selection. Lu et al. (2023) utilizes proprietary chatGPT to tag the instruction data to ensure diversity and complexity. Chen et al. (2023b) utilizes proprietary LLMs chatGPT and Claude2 to assess the quality of the instruction data, generating both ratings and explanations. Du et al. (2023) and Bukharin and Zhao (2023) utilize an extra reward model to assess the quality of data and utilize these scores as a part of their method. Li et al. (2023b) firstly proposes a self-guided method in which no extra LLMs are utilized but still needs to calculate Instruction-Following Difficulty (IFD) scores based on the original pre-trained LLM. Though effective, these methods overly rely on large language models and are too time-consuming to put into practical use.

3 Small Model Proxies for Large Models

The use of proxy models is increasingly recognized in machine learning, particularly when resources are constrained or there is a limited understanding of the original model’s architecture. Chen et al. (2023a) and Hase et al. (2020) demonstrate the utility of lightweight proxy models in evaluating free-text rationales. Similarly, Puigcerver et al. (2021) leverages embeddings from expert models with a k-nearest neighbors classifier to simplify the training of more complex systems. Coleman et al. (2020) and the FAMIE Nguyen et al. (2022) apply downscaled proxy models in fields like image classification and information extraction, utilizing techniques such as layer removal and knowledge distillation for aligning these proxies with larger models. Building on this, Burns et al. (2023) explores the concept of enhancing larger models through weak supervision, training them on labels generated by weaker models. By extending the "Weak to Strong" concept to LLM instruction tuning data selection, our research employs pre-trained smaller models as proxies, demonstrating their effectiveness in assessing instruction complexity, thereby bridging the gap between the comprehensive capabilities of larger models and the agility of smaller ones.

Conclusion

This paper presented “Superfiltering”, a novel and efficient approach for data filtering in the instruction tuning of LLMs. By effectively utilizing weaker models as proxies for evaluating instructional data, particularly in the context of IFD scores, we achieved a significant leap in efficiency, accelerating the data filtering process largely. The experimental results affirm that our method considerably reduces computational overhead while maintaining or even improving the instructional capabilities of LLMs. Thus, Superfiltering stands as a testament to our initial hypothesis and objectives, marking a substantial contribution to the field of natural language processing by offering a scalable, resource-efficient, and effective strategy for the advancement of AI technologies.

References