An agent deploys to resolve setup issues on the codebase, run a minimal reproduction, and estimate full replication cost. Read more below
2
1/N Model capabilities today are subpar for end-to-end autonomous research, what we would categorize as the ability to make “novel discoveries”. However, in our experimentation, research agents have proven to be an excellent way to resolve implementation issues and carry out reproductions of academic works.
3
2/N In this case the objective that is hillclimbed is a 0/1 on whether a paper’s initial claims can be matched. In addition to making previous research codebases reproducible, we also believe a tool like this can make it easy for authors to ensure their codebases are legible in the first place.
4
3/N As models improve, this testbed can extend to more complex tasks that involve actually iterating and improving upon an existing work. Having a public, observable repository of experiments that research agents carry out is going to be invaluable. A slop-free environment where researchers can see exactly what attempts an agent has made when either trying to reproduce a work or make an improvement.