DataSmith
Synthesizing Multilingual Instruction-Tuning Data with the Help of External Tools
Synthesizing supervised finetuning (SFT) data for training multilingual language models is common, but this approach relies heavily on the parametric knowledge of the generator model. This reliance is especially limiting for less-resourced languages, where such knowledge is often sparse or unreliable, risking low-quality data and underperforming models. In this work, we introduce DataSmith, a multilingual data generation framework that reduces this reliance by grounding a tool-use teacher LM with external sources that access the web, reference linguistic knowledge, and consult expert systems. Across 11 languages, this produces quality data for training multilingual models that perform well on downstream cultural, commonsense, and knowledge-based tasks, with the largest gains for underserved languages like Tagalog (+9.3), Czech (+8.3), and Yoruba (+6.8).
Read the full abstract
Synthesizing supervised finetuning (SFT) data for training multilingual language models (LMs) is common, but this approach relies heavily on the parametric knowledge of the generator model. This reliance is especially limiting for less-resourced languages, where such knowledge is often sparse or unreliable, risking low-quality data and underperforming LMs. In this work, we introduce DataSmith, a multilingual data generation framework that reduces this reliance by grounding a tool-use teacher LM with external sources. We equip this LM with tools that can access the web, reference linguistic knowledge, and consult other expert systems. Our experiments across 11 languages (5 high resource, 6 low- to mid-resource) show that DataSmith is capable of producing quality data which can then be used to train multilingual LMs that perform well on downstream cultural knowledge, commonsense reasoning, and knowledge-based tasks. Further analysis of the teacher LM's trajectory reveals that DataSmith invokes more tool-use calls and leans more on linguistic resources for less-resourced languages. Ultimately, we demonstrate through the DataSmith framework that moving away from a model-only paradigm, coupled with a thoughtful curation of external tools, can lead to quality synthetic data for building the next generation of inclusive language technologies.
How DataSmith works
Starting from a set of English seed prompts, DataSmith runs each one through two phases. In the Research Phase, a tool-calling teacher LM decides what it needs to know and gathers grounded evidence by calling web, linguistic, and expert tools, building up a research trajectory. In the Generate Phase, it writes a multi-turn conversation in the target language, conditioned on the grounding information it collected during the research phase. A light post-hoc filter drops malformed or off-language outputs. We will release the generated conversations and research trajectories soon.
See it in action
We generate data across 11 languages (5 high-resource and 6 low- to mid-resource). We show an example of a generated conversation (and its accompanying research phase) below. The Research Phase is the teacher's tool-use trajectory (its thinking, tool calls, and the results it read). The Generate Phase is the multi-turn conversation it then produced.
Research Phase
Generate Phase
Does DataSmith produce better supervision data?
SetupIn order to test whether DataSmith is a better alternative for multilingual synthetic data generation, we measure dataset quality by training a student model on each synthetic dataset and evaluating it on downstream tasks.
- Target languages: We test 11 languages (5 high-resource and 6 low- to mid-resource): Arabic, German, French, Spanish, Japanese, Czech, Indonesian, Tagalog, Cebuano, Swahili, and Yoruba.
- Pipeline setup: We use Kimi-K2.5 as the tool-use teacher LM and generate 5k multi-turn dialogues per language. For the main experiments, we run continuous SFT on Gemma 3 12B Instruct as the shared base student model.
- Benchmarks: We evaluate on Global MMLU (factual knowledge), Global PIQA (physical reasoning), and CulturalBench (cultural knowledge), and report per-language averages over three runs.
- Baselines: We compare against Teacher-only (same teacher LM without tools) and Translate-Train (5k English seed conversations translated with TranslateGemma 27B). All methods use the same base model and hyperparameters for a fair comparison.
ResultsDataSmith improves over the base model and outperforms both baselines, across languages and across benchmarks. It achieves the best score in 10 of 11 languages, trailing only on French (-0.5pp), and the largest gains come from low- to mid-resource languages such as Tagalog (+9.3pp), Czech (+8.3pp), Yoruba (+6.8pp), and Cebuano (+6.6pp). It also ranks first on all three benchmarks, with the largest improvement on Global MMLU (45.5 to 56.9, +11.4pp). These findings generalize to other base model families, including Tiny Aya Global 3B and Llama 3.1.
| Method | Avg | High-resource | Low- to mid-resource | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ara | deu | fra | spa | jpn | ces | ind | tgl | ceb | swa | yor | ||
| Base (no SFT) | 63.2 | 66.1 | 67.9 | 64.0 | 72.0 | 71.2 | 57.7 | 70.9 | 55.4 | 45.6 | 71.3 | 53.4 |
| + DataSmith (ours) | 67.9 | 68.3 | 70.2 | 70.3 | 74.6 | 73.8 | 66.0 | 74.8 | 64.7 | 52.2 | 72.3 | 60.2 |
| + Teacher-only | 65.4 | 67.7 | 68.3 | 70.8 | 71.8 | 72.7 | 64.1 | 72.5 | 57.7 | 48.9 | 68.2 | 56.8 |
| + Translate-Train | 55.5 | 53.2 | 62.3 | 63.6 | 62.7 | 60.1 | 55.0 | 67.0 | 55.9 | 31.1 | 62.0 | 37.4 |
| Method | Global MMLU | Global PIQA | CulturalBench | ||||||
|---|---|---|---|---|---|---|---|---|---|
| High | Low-Mid | Avg | High | Low-Mid | Avg | High | Low-Mid | Avg | |
| Base (no SFT) | 60.7 | 45.5 | 53.1 | 81.6 | 79.2 | 80.4 | 61.5 | 55.5 | 58.5 |
| + DataSmith (ours) | 65.3 | 56.9 | 61.1 | 82.8 | 82.4 | 82.6 | 65.7 | 59.5 | 62.6 |
| + Teacher-only | 65.3 | 52.9 | 59.1 | 80.7 | 77.6 | 79.2 | 64.0 | 57.2 | 60.6 |
| + Translate-Train | 57.7 | 53.7 | 55.7 | 73.8 | 68.6 | 71.2 | 48.8 | 37.9 | 43.4 |
Analysis of research trajectories
Experiment: How does tool use vary across languages?
SetupWe track DataSmith's research trajectory, i.e., the sequence of tool calls it issues before writing each dialogue, and assign every call to one of three families: web, linguistic, and expert. We then average the number of calls per dialogue within each resource tier.
ResultsDataSmith issues more tool calls for low- to mid-resource languages (5.0 to 6.7 on average) than for high-resource ones (4.5 to 5.1), and a larger fraction of those calls go to linguistic resources. This suggests that low-resource languages need grounding in the language itself, such as vocabulary and usage, and not just in factual knowledge.
Tool use by language. Average number of tool-use calls per dialogue, broken down by tool category, for high-resource (left of the dotted line) and low- to mid-resource (right) languages. Low- to mid-resource languages invoke more tools overall and rely more on linguistic resources, while web tools dominate across all languages.
Tool use across the research stage. The teacher LM's action at each step of the research stage, with <|end_research|> marking turns where it is satisfied and ends research. The teacher leans heavily on web search in early turns. For low- to mid-resource languages, it researches for longer and relies more on linguistic and expert tools.
Experiment: How good is DataSmith's research stage?
Quality of DataSmith's research trajectories across languages (min 1.0, max 5.0). We report the mean LLM-as-a-judge score (GPT-OSS 120B) over 1k trajectories per language along three axes.
SetupWe score each trajectory with an LLM judge along three rubrics (1 to 5): evidence groundedness (is the dialogue supported by what the teacher actually retrieved), information density (how much of the retrieved information is useful), and research efficiency (does it avoid redundant calls). We sample 1k trajectories per language over three trials. In order to check that the judge is reliable, we compare its scores against a human annotator. We find a weighted Cohen's kappa of 0.769, which indicates high agreement.
ResultsTrajectories are consistently well-grounded in evidence across languages (around 4.1), meaning the generated conversations rely on gathered information rather than the teacher's priors. Research efficiency is the weakest rubric, and the drop is more pronounced for low-resource languages such as Yoruba and Swahili. Overall scores range from 3.35 to 3.77.
Conclusion
In this work, we introduced DataSmith, a multilingual data generation framework that grounds a teacher LM in external knowledge through tool-use, drawing on the web, linguistic resources, and expert systems. Across 11 languages and three benchmarks, student models trained on DataSmith-generated data outperform those trained with a classic teacher-only pipeline and a translate-train setup, with the largest gains on low- to mid-resource languages. Our analysis of the teacher's research trajectories shows that this grounding is adaptive: the teacher issues more tool calls and leans more heavily on linguistic resources specifically for the languages where its parametric knowledge is weakest. Together, these results suggest a path to better multilingual supervision data: equipping teachers with access to the right external sources. We hope DataSmith paves the way toward more equitable language technologies for underserved languages.
Limitations
Our work comes with limitations. First, the quality of the data generated by DataSmith depends on the design of its harness and the choice of external tools. We mitigate this by grounding our tool choice in sources commonly used for constructing multilingual datasets, but alternative tool combinations may yield different results, which we leave for future work. Second, our framework assumes that these tools adequately cover the target language. This assumption weakens as we move down the resource hierarchy, as Wikipedia coverage varies widely in both size and quality across languages and many languages lack digitized grammar books or lexicons. In addition, expert translation and language identification models only support a subset of the world's languages. Finally, our experiments use a teacher LM with strong function-calling capabilities over long horizons, such as Kimi-K2.5. Weaker teachers may fail to execute the research pipeline reliably, and we have not characterized how data quality degrades as the teacher's tool-use ability decreases.
Ethics Statement
Synthetic data generation carries the risk of perpetuating biases present in the models involved in the pipeline. Although DataSmith grounds generation in external sources, the teacher LM, the expert systems it consults, and even sources such as web search results and Wikipedia can encode societal and cultural biases, which may propagate into the generated conversations and, ultimately, into the student models trained on them. We partially mitigate this through grounding, but we did not perform a dedicated audit of the generated datasets, and we encourage practitioners to apply appropriate filtering and auditing before deployment.
Citation
@techreport{miranda2026datasmith,
title = {DataSmith: Synthesizing Multilingual Instruction-Tuning Data with the Help of External Tools},
author = {Miranda, Lester James V. and Korhonen, Anna},
year = {2026},
month = jul,
note = {v1 posted on July 6, 2026},
institution = {Language Technology Lab, University of Cambridge},
type = {Technical Report},
url = {https://github.com/ljvmiranda921/datasmith/blob/main/docs/DataSmith-TechReport.pdf}
}