University of Cambridge Language Technology Lab

DataSmith

Synthesizing Multilingual Instruction-Tuning Data with the Help of External Tools

Lester James V. Miranda, Anna Korhonen
Language Technology Lab, University of Cambridge
Correspondence to ljvm2@cam.ac.uk
Status: In Progress (v1 posted on July 6, 2026)

Synthesizing supervised finetuning (SFT) data for training multilingual language models is common, but this approach relies heavily on the parametric knowledge of the generator model. This reliance is especially limiting for less-resourced languages, where such knowledge is often sparse or unreliable, risking low-quality data and underperforming models. In this work, we introduce DataSmith, a multilingual data generation framework that reduces this reliance by grounding a tool-use teacher LM with external sources that access the web, reference linguistic knowledge, and consult expert systems. Across 11 languages, this produces quality data for training multilingual models that perform well on downstream cultural, commonsense, and knowledge-based tasks, with the largest gains for underserved languages like Tagalog (+9.3), Czech (+8.3), and Yoruba (+6.8).

Read the full abstract

Synthesizing supervised finetuning (SFT) data for training multilingual language models (LMs) is common, but this approach relies heavily on the parametric knowledge of the generator model. This reliance is especially limiting for less-resourced languages, where such knowledge is often sparse or unreliable, risking low-quality data and underperforming LMs. In this work, we introduce DataSmith, a multilingual data generation framework that reduces this reliance by grounding a tool-use teacher LM with external sources. We equip this LM with tools that can access the web, reference linguistic knowledge, and consult other expert systems. Our experiments across 11 languages (5 high resource, 6 low- to mid-resource) show that DataSmith is capable of producing quality data which can then be used to train multilingual LMs that perform well on downstream cultural knowledge, commonsense reasoning, and knowledge-based tasks. Further analysis of the teacher LM's trajectory reveals that DataSmith invokes more tool-use calls and leans more on linguistic resources for less-resourced languages. Ultimately, we demonstrate through the DataSmith framework that moving away from a model-only paradigm, coupled with a thoughtful curation of external tools, can lead to quality synthetic data for building the next generation of inclusive language technologies.

How DataSmith works

Starting from a set of English seed prompts, DataSmith runs each one through two phases. In the Research Phase, a tool-calling teacher LM decides what it needs to know and gathers grounded evidence by calling web, linguistic, and expert tools, building up a research trajectory. In the Generate Phase, it writes a multi-turn conversation in the target language, conditioned on the grounding information it collected during the research phase. A light post-hoc filter drops malformed or off-language outputs. We will release the generated conversations and research trajectories soon.

The DataSmith framework: a tool-calling teacher LM grounded in web, linguistic, and expert tools, running a research phase then a generation phase.
Overview of the DataSmith framework and its comparison to a standard synthetic data pipeline. While a classic synthetic pipeline is restricted to its own parametric knowledge when generating data, DataSmith is augmented with external tools such as web search, linguistic resources, and expert systems, grounding the generation process in external knowledge.

See it in action

We generate data across 11 languages (5 high-resource and 6 low- to mid-resource). We show an example of a generated conversation (and its accompanying research phase) below. The Research Phase is the teacher's tool-use trajectory (its thinking, tool calls, and the results it read). The Generate Phase is the multi-turn conversation it then produced.

High-resource
Low- to mid-resource
Seed prompt:

Research Phase

Generate Phase

Does DataSmith produce better supervision data?

In order to test whether DataSmith is a better alternative for multilingual synthetic data generation, we measure dataset quality by training a student model on each synthetic dataset and evaluating it on downstream tasks.

DataSmith improves over the base model and outperforms both baselines, across languages and across benchmarks. It achieves the best score in 10 of 11 languages, trailing only on French (-0.5pp), and the largest gains come from low- to mid-resource languages such as Tagalog (+9.3pp), Czech (+8.3pp), Yoruba (+6.8pp), and Cebuano (+6.6pp). It also ranks first on all three benchmarks, with the largest improvement on Global MMLU (45.5 to 56.9, +11.4pp). These findings generalize to other base model families, including Tiny Aya Global 3B and Llama 3.1.

Main results across all languages. Accuracy of a Gemma 3 12B Instruct model fine-tuned on data from each method, averaged across Global MMLU, Global PIQA, and CulturalBench and over three runs. Avg is the macro-average across all languages.
MethodAvgHigh-resourceLow- to mid-resource
aradeufraspajpncesindtglcebswayor
Base (no SFT)63.266.167.964.072.071.257.770.955.445.671.353.4
+ DataSmith (ours)67.968.370.270.374.673.866.074.864.752.272.360.2
+ Teacher-only65.467.768.370.871.872.764.172.557.748.968.256.8
+ Translate-Train55.553.262.363.662.760.155.067.055.931.162.037.4
Results for each benchmark, broken down by resource level. Downstream performance of a Gemma 3 12B Instruct model fine-tuned on data from each method, averaged over three runs per language. Avg is the macro-average across resource levels.
MethodGlobal MMLUGlobal PIQACulturalBench
HighLow-MidAvgHighLow-MidAvgHighLow-MidAvg
Base (no SFT)60.745.553.181.679.280.461.555.558.5
+ DataSmith (ours)65.356.961.182.882.482.665.759.562.6
+ Teacher-only65.352.959.180.777.679.264.057.260.6
+ Translate-Train57.753.755.773.868.671.248.837.943.4

Analysis of research trajectories

Experiment: How does tool use vary across languages?

We track DataSmith's research trajectory, i.e., the sequence of tool calls it issues before writing each dialogue, and assign every call to one of three families: web, linguistic, and expert. We then average the number of calls per dialogue within each resource tier.

DataSmith issues more tool calls for low- to mid-resource languages (5.0 to 6.7 on average) than for high-resource ones (4.5 to 5.1), and a larger fraction of those calls go to linguistic resources. This suggests that low-resource languages need grounding in the language itself, such as vocabulary and usage, and not just in factual knowledge.

Tool use by language. Average number of tool-use calls per dialogue, broken down by tool category, for high-resource (left of the dotted line) and low- to mid-resource (right) languages. Low- to mid-resource languages invoke more tools overall and rely more on linguistic resources, while web tools dominate across all languages.

Tool use across the research stage. The teacher LM's action at each step of the research stage, with <|end_research|> marking turns where it is satisfied and ends research. The teacher leans heavily on web search in early turns. For low- to mid-resource languages, it researches for longer and relies more on linguistic and expert tools.

Experiment: How good is DataSmith's research stage?

Quality of DataSmith's research trajectories across languages (min 1.0, max 5.0). We report the mean LLM-as-a-judge score (GPT-OSS 120B) over 1k trajectories per language along three axes.

We score each trajectory with an LLM judge along three rubrics (1 to 5): evidence groundedness (is the dialogue supported by what the teacher actually retrieved), information density (how much of the retrieved information is useful), and research efficiency (does it avoid redundant calls). We sample 1k trajectories per language over three trials. In order to check that the judge is reliable, we compare its scores against a human annotator. We find a weighted Cohen's kappa of 0.769, which indicates high agreement.

Trajectories are consistently well-grounded in evidence across languages (around 4.1), meaning the generated conversations rely on gathered information rather than the teacher's priors. Research efficiency is the weakest rubric, and the drop is more pronounced for low-resource languages such as Yoruba and Swahili. Overall scores range from 3.35 to 3.77.

Conclusion

In this work, we introduced DataSmith, a multilingual data generation framework that grounds a teacher LM in external knowledge through tool-use, drawing on the web, linguistic resources, and expert systems. Across 11 languages and three benchmarks, student models trained on DataSmith-generated data outperform those trained with a classic teacher-only pipeline and a translate-train setup, with the largest gains on low- to mid-resource languages. Our analysis of the teacher's research trajectories shows that this grounding is adaptive: the teacher issues more tool calls and leans more heavily on linguistic resources specifically for the languages where its parametric knowledge is weakest. Together, these results suggest a path to better multilingual supervision data: equipping teachers with access to the right external sources. We hope DataSmith paves the way toward more equitable language technologies for underserved languages.

Limitations

Our work comes with limitations. First, the quality of the data generated by DataSmith depends on the design of its harness and the choice of external tools. We mitigate this by grounding our tool choice in sources commonly used for constructing multilingual datasets, but alternative tool combinations may yield different results, which we leave for future work. Second, our framework assumes that these tools adequately cover the target language. This assumption weakens as we move down the resource hierarchy, as Wikipedia coverage varies widely in both size and quality across languages and many languages lack digitized grammar books or lexicons. In addition, expert translation and language identification models only support a subset of the world's languages. Finally, our experiments use a teacher LM with strong function-calling capabilities over long horizons, such as Kimi-K2.5. Weaker teachers may fail to execute the research pipeline reliably, and we have not characterized how data quality degrades as the teacher's tool-use ability decreases.

Ethics Statement

Synthetic data generation carries the risk of perpetuating biases present in the models involved in the pipeline. Although DataSmith grounds generation in external sources, the teacher LM, the expert systems it consults, and even sources such as web search results and Wikipedia can encode societal and cultural biases, which may propagate into the generated conversations and, ultimately, into the student models trained on them. We partially mitigate this through grounding, but we did not perform a dedicated audit of the generated datasets, and we encourage practitioners to apply appropriate filtering and auditing before deployment.

Citation

@techreport{miranda2026datasmith,
  title       = {DataSmith: Synthesizing Multilingual Instruction-Tuning Data with the Help of External Tools},
  author      = {Miranda, Lester James V. and Korhonen, Anna},
  year        = {2026},
  month       = jul,
  note        = {v1 posted on July 6, 2026},
  institution = {Language Technology Lab, University of Cambridge},
  type        = {Technical Report},
  url         = {https://github.com/ljvmiranda921/datasmith/blob/main/docs/DataSmith-TechReport.pdf}
}