1

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI
We introduce MLS-Bench, a benchmark of 140 tasks across 12 ML domains that evaluates whether AI systems can invent generalizable and scalable ML methods, and find that current agents remain far from reliably surpassing human-designed methods.
ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving
Stealthy Backdoor Attack in Federated Learning via Adaptive Layer-wise Gradient Alignment