Had a sort of a funny experience with OpenAI’s Deep Research tool,...

@littmath
Daniel Litt@littmath
14 views Feb 19, 2025 ~2 min read
1
Had a sort of a funny experience with OpenAI’s Deep Research tool, which I wanted to share since I think it reveals some of the tool’s strengths and weaknesses.
2
.@srivatsamath recently suggested to me that (as a result of the Internet democratizing access to advanced math), there’s been an increase in important math research done by young people. I was curious if this is true.
3
This seemed to me to be a good use case for Deep Research. As a (admittedly poor) proxy for the question, I asked it about the age of authors publishing in the Annals of Mathematics, arguably one of the top math journals, from 1950-2025.
4
This is a pretty involved request—Annals has published ~3000 papers in this time, and one has to look up each of the authors and figure out/estimate their birth date. Doing this would be tedious but trivial for a human. I thought it would be within the purview of Deep Research.
5
The tool produced a beautifully written and argued report—and contrary to the original suggestion, it seemed to be saying average age of publication in the Annals has been *increasing* with time. Here are some summary statistics it produced.
Media image
6
It also performed a few other analyses—looking at age at the time of first publication, for example.
7
The only problem is that it’s all made up. Despite claiming to have looked at every paper published in the Annals in this 75-year period, poking around the pages it looked at suggests it only looked at ~5-6 papers.
Media image
8
Reconstructing the tool’s approach, it seems to have run across an article challenging Hardy’s claim that math research is a “young man’s [sic] game,” and backfilled a narrative to support this challenge.
Media image
9
I made a further request, which was just for it to generate a dataset of authors, Annals publications, and author’s age at time of publication. It produced this:
Media image
10
In other words, it claimed again to produce a complete dataset but in fact only produced ~7 lines, with a placeholder for the other ~3000.
11
I think in retrospect I should have expected these failures—going through ~3000 papers would take many days for a human, and AI tools just don’t have this capacity yet. But the failure mode was very far from graceful!
12
I hope/expect we will see tools able to do this kind of thing in the near- or medium-term. But I have to wonder how many people are treating these tools as having this kind of capability currently.
13
Anyway, you can see the full reports/prompts here: chatgpt.com/share/67b4a32b…
Actions
What You Can Do
  • Export as PDF or Markdown
  • Batch Export to Notion
  • Bookmark & Highlight
  • LinkedIn & Instagram Carousel Maker
Create Free Account

Includes 7-day Premium trial