Research
Bootstrapping Code Translation with Weighted Multilanguage Exploration
arXiv:2601.03512v2 Announce Type: replace-cross Abstract: Code translation across multiple programming languages is essential yet challenging due to two vital obstacles: scarcity of parallel data pair
arXiv:2601.03512v2 Announce Type: replace-cross Abstract: Code translation across multiple programming languages is essential yet challenging due to two vital obstacles: scarcity of parallel data paired with executable test oracles, and optimization imbalance when handling diverse language pairs. We propose BootTrans, a bootstrapping method that resolves both obstacles. Its key idea is to leverage the functional invariance and cross-lingual portability of test suites, adapting abundant pivot-language unit tests to serve as universal verification oracles for multilingual reinforcement learning (RL) training. Our method introduces a dual-pool architecture with seed and exploration pools to progressively expand training data via execution-guided experience collection. Furthermore, we design a language-aware weighting mechanism that dynamically prioritizes harder translation directions based on relative performance across sibling languages, mitigating optimization imbalance. Extensive experiments on the HumanEval-X and TransCoder-Test benchmarks demonstrate substantial improvements over baseline LLMs across all translation directions, with ablation studies validating the effectiveness of both bootstrapping and weighting components.
Related
- Fine-Tuning Code Language Models to Detect Cross-Language Bugs
- Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction
- Should We be Pedantic About Reasoning Errors in Machine Translation?
- DeepGuard: Secure Code Generation via Multi-Layer Semantic Aggregation
Source: arXiv cs.AI | 2026-04-22