Jiefeng Chen, Bhavana Dalvi Mishra, Jaehyun Nam, Rui Meng, Tomas Pfister, Jinsung Yoon
5 min
Automating machine learning engineering (MLE) is significantly more complex than general software engineering because it requires navigating resource-intensive model training, opaque performance attribution, and high architectural complexity. Existing LLM-based agents often fail to manage these constraints, producing monolithic, fragile scripts that ignore the economic reality of computational budgets and lack the ability to learn from past experimental failures.
MARS (Modular Agent with Reflective Search) addresses these bottlenecks by reformulating research as a search for an optimal software repository. The framework is built on three core pillars:
Evaluated on the MLE-Bench benchmark, MARS establishes a new state-of-the-art among open-source frameworks. It consistently outperforms existing agents across various task complexities, demonstrating a superior ability to navigate long-horizon exploration. Qualitative analysis reveals that 63% of the agent's strategic lessons originate from cross-branch transfer, confirming that MARS effectively generalizes insights across different search paths. By mimicking the structured, iterative approach of human engineers, MARS provides a scalable solution for autonomous scientific discovery.
A critical bottleneck in automating AI research is the execution of complex machine learning engineering (MLE) tasks. MLE differs from general software engineering due to computationally expensive evaluation (e.g., model training) and opaque performance attribution. Current LLM-based agents struggle here, often generating monolithic scripts that ignore execution costs and causal factors. We introduce MARS (Modular Agent with Reflective Search), a framework optimized for autonomous AI research. MARS relies on three pillars: (1) Budget-Aware Planning via cost-constrained Monte Carlo Tree Search (MCTS) to explicitly balance performance with execution expense; (2) Modular Construction, employing a "Design-Decompose-Implement" pipeline to manage complex research repositories; and (3) Comparative Reflective Memory, which addresses credit assignment by analyzing solution differences to distill high-signal insights. MARS achieves state-of-the-art performance among open-source frameworks on MLE-Bench under comparable settings, maintaining competitiveness with the global leaderboard's top methods. Furthermore, the system exhibits qualitative "Aha!" moments, where 63% of all utilized lessons originate from cross-branch transfer, demonstrating that the agent effectively generalizes insights across search paths.
Sam: Right. It moves away from "generate-and-pray" debugging toward localizing where behavior deviates from expectation. And it matters for the memory too. If you know which module changed, you can tie a performance delta to that block, so the memory records the history of modifications rather than just similar code. The idea is that it learns patterns of modular modification, a research-engineering strategy rather than syntax.
Alex: This is where I'd push back as a referee. The memory is described as causal, but the failure analysis I've seen calculates correlations between error magnitude and features like label cardinality. Is that enough to say the agent understands cause?
Sam: It isn't, and that's a critical distinction. It's a heuristic, a diagnostic signal rather than proof. If error scales with label cardinality, the agent infers the model struggles with high-density targets. That doesn't establish the flaw, but it gives a concrete direction for the next ablation.
Alex: So it's a proxy for causality. It tells the agent where to look, not why the model failed. "Causal history" oversells what is closer to a sophisticated error log.
Sam: Fair. The attribution is LLM-based, so it's susceptible to hallucination. The agent could credit a gain to a code change that was really stochastic noise. This is the main limitation: the system is only as good as the agent's reasoning about the execution trace.
Alex: And if the attribution is wrong, the distilled lesson is a poisoned data point. Do they check these attributions beyond asking the LLM to summarize the diff?
Sam: There's no formal verification of the attributions yet. Unit tests check that modules behave as specified, not that the explanation for a gain is right. The authors point to formal program analysis as future work. For now, modularity is the mitigation, because it bounds how far a wrong attribution can do damage.
Alex: There's a second boundary too. Does the modular design hurt when a problem doesn't fit a standard pipeline?
Sam: It likely does. MARS suits tasks with a logical engineering flow. A problem that needs a novel, non-linear architecture resisting standard decomposition could trip it up. It's a tool for systematic exploration, not a substitute for creative leaps.
Alex: A pragmatic trade-off, then. If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Sam: Thanks for listening.