Model Releases

🚨 Are we witnessing the automation of AI research? @HuggingFace just unveiled 'ML-Intern' and my mind is BLOWN 🤯 It’s an open-source pipel…

🚨 Are we witnessing the automation of AI research? @HuggingFace just unveiled 'ML-Intern' and my mind is BLOWN 🤯 It’s an open-source pipeline that replicates the exact daily loop of an ML researcher.

DGX agentx-post
model-releasesclem-delangue--x

🚨 Are we witnessing the automation of AI research? @HuggingFace just unveiled "ML-Intern" and my mind is BLOWN 🤯 It’s an open-source pipeline that replicates the exact daily loop of an ML researcher. You simply write a prompt, then watch the magic happen: → ML-Intern reads the arXiv papers → digs through citations → spins up GPU sandboxes → iterates → ... even builds you a deeply researched model Awesome, right? Plus, everything happens right inside the HF ecosystem: šŸ”¹ training via HF Jobs šŸ”¹ monitoring via Trackio šŸ”¹ dataset loading & final model pushing straight to the Hub The proof is in the pudding! They let ML-Intern tackle scientific reasoning. That thing autonomously: - researched official benchmarks - found OpenScience and NemoTron-CrossThink - pulled 7 difficulty-filtered variants (ARC/SciQ/MMLU) - ... and ran 12 SFT runs on Qwen3-1.7B The results? GPQA jumped from 10% to 32% in under 10 hours! As a comparison, Claude Code topped out at 22.99%. This secured SOTA on PostTrainBench, the rigorous new benchmark by ELLIS Institute Tübingen and Max Planck Institute for 10-hour LLM post-training. It’s out now, and completely free and open-source. You've got to dig into this code and app below 🧵 ↓ Media

Related

Source: Clem Delangue (X) | 2026-04-22

Loading related sources…