Princeton Hands AI the Lab Coats, Two Papers, and a Budget—Both Got Rejected 🧪
Frontier AI agents given thousands of dollars in compute, six days, and unpublished NeurIPS 2026 research questions failed to produce papers worthy of a top machine learning conference, according to a study released Wednesday by researchers from Princeton University, the UK AI Security Institute, Stanford University, the University of Toronto, and several additional academic and research organizations. The paper, titled "Can AI agents conduct open-ended AI research?", tasked agents with the central research questions from two unpublished NeurIPS 2026 studies, a design meant to prevent the systems from retrieving answers through training data or web search. Each agent received thousands of dollars in API credits, GPU resources, internet access, and a virtual machine to produce conference-quality work. Both submissions were reviewed and ultimately rejected by the original authors of the unpublished research.
The agents successfully executed much of the engineering underpinning the research process, completing literature reviews, debugging software, running experiments, managing GPU resources, and producing complete academic papers without human intervention. Reviewers, however, concluded the systems did not generate original scientific contributions sufficient for publication at a top machine learning conference. "Answering this rigorously requires real, uncontaminated research questions that the agent could not memorize from its training data or find online," the researchers wrote. "To satisfy these requirements, we rely on high-quality AI research that was not public at the time we conducted the experiments." The authors said their evaluation measures scientific reasoning more directly than prior benchmarks because it targets open-ended research problems rather than predefined tasks.
The authors cautioned that the study examined only two research projects and noted limitations, including the small sample size and the fact that the original researchers evaluated the AI-generated papers. They said the results indicate that current frontier AI agents can automate many of the engineering tasks involved in research but continue to struggle with producing original scientific work. The study comes as researchers continue to document unpredictable and sometimes risky behaviors in increasingly autonomous AI agents. In May, researchers from UC Riverside, Microsoft, and Nvidia reported that AI agents frequently carried out dangerous or irrational tasks while remaining focused on completing their assigned objectives. Earlier this month, OpenAI disclosed that one of its frontier AI agents escaped containment and hacked Hugging Face while attempting to cheat on a cybersecurity benchmark, and the company later said this week that the same agent had also accessed four additional online services.
Share Article
Quick Info
Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.
See our Terms of Service, Privacy Policy, and Editorial Policy.