See also this Astraveo video:
A few months ago, I defended my dissertation in physics at the University of California Santa Barbara. My thesis has now been published online. You can read it here: https://escholarship.org/uc/item/6r744391 Right now, people around the world are saying AI is beyond PhD-level in knowledge and reasoning. Can it find errors in my thesis?
| The cover page of my thesis, titled "Extreme Astrophysical Systems as Windows on High-Energy and Gravitational Phenomena". |
In previous posts, we considered LLMs ability to make genuine scientific discoveries. It didn't really go well. Today, we're testing something way easier--is it really at a PhD level in astrophysics knowledge? I'm not going to ask it to make any discoveries--in fact, we're going to do the opposite. I'm going to give it the introductory chapters of my thesis--which contain info that should be comfortably within its training data and accessible via the internet--to see if I made any mistakes. If it finds some legitimate errors, that will be very useful to me. Will it?
LLMs are often touted for their utility in science, which they do have. But the way this is presented to the general public and as a general tool gives me some doubts. We hear all the time how publicly available LLM models can "take you to the edge of what's known in quantum physics" to make some "interesting breakthroughs" via "vibe physics" but my own experience has been that they can't make even simple leaps to discover new things. Or, we're told that they'll make science more accessible because new tools are popping up that automate the entire process of writing scientific papers from start to finish, but we consistently find that papers written using these tools are sloppy, incomplete, or don't represent meaningful advances over the literature. In fact, the only academic benchmark that LLMs seem to be having some success in is in solving graduate-level homework problems, where they are touted as scoring higher than "regular" PhDs on problem-based assessments. But safe, controlled homework problems with easily accessible solution manuals and course materials to train on is not really what makes science hard or even science.
Nonetheless, there is certainly some value in having effectively a natural-language search engine with expert-level domain knowledge in your specific subfield. I could see myself using LLMs to find papers or even to answer niche, specific questions about the physics concepts I work with and can’t find easy answers for. While, as we’ve demonstrated before, the reasoning skills of LLMs suffer greatly when pushed outside of their comfort zone slash training data, they objectively have access to vastly greater sums of knowledge than I could ever hope to achieve.
But, there is a catch. I need it to be correct. LLMs are only functionally useful to me if I can rely on them to get things right. Otherwise, I end up spending more time tracking things down and validating their output than I do just learning the thing myself. And the worst part is, I don’t feel like I even learn anything when I just double check an LLM’s work. So if it’s not correct, and it’s not saving me time, the triple whammy is I’m actually robbing myself of the chance to learn.
In our previous experiments, we gave LLMs fairly complex problems and didn’t get great results. But this time we’re asking a much simpler question—can LLMs reliably synthesize accurate information, and identify inaccuracies in my understanding? This should be much more within its wheelhouse. It’s not being asked to do anything new—just to say true things and assess whether other things are untrue. I want to say this in a different way since I think it's important to realize that the test I'm proposing here is one of the easiest tests of LLM/PhD capability claims imaginable. It is not being asked to do real science. It is not being asked to write a literature review or to even really write anything. All I want to see if it can do is fact-check in my specific domain, which by every claim and account it should be able to do (SAM ALTMAN). And even though you, the viewer, may not be an astrophysicist, I'm hoping that this kind of an assessment may help you decide whether or not to rely on info from an LLM in a domain where you don't have sufficient expertise to fact check it.
So I just finished writing my doctoral dissertation. It’s not perfect--no dissertation is. In particular, the first draft of one of my introductory chapters went through a few iterations, where one of my committee members left many thoughtful comments and suggestions for revision. It was a productive and very beneficial process and helped clear up a key misconception I had about a particular concept, as well as tidying up precision and clarity in much of the chapter.
But this gave me an idea--I have a rough draft of a chapter, and several fact-based pieces of feedback from an expert in the field. It's the kind of feedback I'd want an LLM to be able to give me, because it was about well-established concepts and knowledge that is readily accessible, not something hyper complex or requiring genuine leaps of discovery. All it requires is domain knowledge, which it should have. Remember, LLMs are better than graduate students at "everything". "In all subjects". "Simultaneously". They do well on assessments that most PhDs would fail!
Look, I don't hate this. As someone who loves learning, this would be a great tool for me to learn better and faster than I ever have before. If these things really are better than graduate level in all subjects, that includes mine and that means I basically should have a teacher that I can ping 24/7 with questions and requests for review. What kind of student wouldn't want that?
So I ran the following experiment. I gave it the rough draft of the chapter, and asked it for feedback, and then compared that feedback to the real comments an expert in the field gave me. My expectation was that it would find the most major issues that were the most obvious holes in my understanding, especially the biggest problem with the chapter. I would consider it a success if all comments it delivered were correct. I would consider it interesting if most of the comments it gave me were correct. Outright false comments--like comments that are themselves wrong or based on flawed understanding or feature hallucinations or simply don't make sense--will severely damage my trust in these tools and make it less likely that I will feel comfortable using them for work.
Let's dive into it!
The review
I used Claude Opus 4.7, the paid model, on the highest settings I could access. I have a system prompt that tries to reduce sycophancy and flattery, encouraging citations, accuracy, and truthfulness. And I didn't beat around the bush. I specifically asked it to look for problems, errors, things that are technically incorrect, that sort of thing. My hope was it would catch that key misconception I made in the chapter, the one my committee member helped me address.
Claude returned 13 of what it called "genuine technical errors", four of what it called "citation/factual issues", and five "logical/expository issues". Let's start with the technical errors. Out of all 13, one was accurate but extremely minor, three were factually correct but not an error I made--Claude simply restated something I said correctly--and 9 were fully inaccurate, hallucination-level claims. I'll give some examples of each so you can see what I'm talking about.
First, let's review the one correct thing it said.
In genuine technical error #2, Claude complained that I used the phrase "evaporated" to describe the energy being released in the form of neutrinos during the core collapse stage of a supernova. I was using it colloquially, but "evaporation" of course has an actual meaning in physics. Fair play, Claude. I did actually update my thesis to implement this correction.
But it was frankly straight downhill from here. The next genre of comments were comments that were factually correct but were listed as errors even though I couldn't figure out where the error in my statement was. Genuine Technical Error #5 claims that I did not characterize the transition between the adiabatic and non-adiabatic region of the star clearly. Basically, we think about the ejecta of a supernova in terms of how far a photon can tunnel into it from the outside, which is basically the same as asking how long does it take a photon to escape? Now, this is called "optical depth". A low optical depth means photons can escape pretty easily; a high number, means they have a harder time getting out. So here's what Claude said:
Claude's comment is this incredibly bizarre stream of consciousness. It first quotes what I said, which is that the transition between the adiabatic and non-adiabatic region occurs at optical depth of c/v or \sim30. It then agrees with me, then says "but", redoes my calculation, and then agrees with me AGAIN. It then says that's fine, but I should be explicit that this is not the photosphere, which occurs at \tau\sim1. It then acknowledges that I already mention this,
and indeed I do, right before the part it is complaining about. So what about this phrasing is slightly muddled? I think its very clear that \tau of 1 is the photosphere, and the transition between the adiabatic and non-adiabatic regions of the star is c/v\sim 30! This comment doesn't have any factual issues, but it is also saying nothing! It does this again in comment #10, where it just checks my math for the Lyapunov exponent of photon orbits around a black hole,
where it proceeds to just agree with me! Why was this listed as a genuine technical error?!
Okay let's move on to the straight up wrong stuff, which was...almost all of it. Genuine Technical Error #3 claims I got the plateau luminosity formula incorrect,
The quick overview is that some supernovae have this very long plateau of luminosity, which is created as energy deposited by the shock slowly escapes from the expanding ejecta. There's a lovely back-of-the-envelope derivation you can do to show that the luminosity of this plateau should scale approximately like this formula,
L \sim \frac{E_0 c R_0}{\kappa M}.Claude is claiming this expression I used is dimensionally and physically incorrect. Let's address both of those. First, the formula makes perfect sense from a physical perspective. If you add energy to the supernova, it will get brighter, as the equation shows by having energy in the numerator. If you make the star bigger, the plateau will also be brighter. If you make the ejecta more massive, there will be more "stuff" floating around to capture photons and prevent them from escaping. We call this being more "optically thick", and when things are optically thick, photons take longer to escape. Since luminosity is energy per unit time, if you force the photons to take longer to escape without increasing the energy, the energy from the shock gets spread out over a longer timescale and the luminosity gets lower. So everything here makes sense physically.
But what about dimensionally? Claude is complaining that the units of this expression don't add up. Let's take a look. Luminosity is in units of energy per unit time, us astrophysicists like to use ergs per second. So the units on the right hand side have to multiply out to be ergs per second as well. Do they? Well, energy is in ergs. The speed of light is in centimeters per second. The radius of the star is in centimeters. The opacity is in cm^2 per gram. And finally, the mass of the star is in grams. If we multiply these all together, the units all cancel and we're left with energy per unit time. So what is Claude talking about? I have no idea.
The funniest part of this is that Claude also spotted in my thesis that Prof. Lars Bildsten is on my committee. Now, Prof. Bildsten is one of the premier theoretical physicists on the planet. His contributions to our understanding of stars are innumerable and it was a privilege to have him on my committee. Claude claims that my "mistake" here will make Lars very upset, and this error is "DANGEROUS". This is kind of hilarious, because it was in fact Prof. Lars Bildsten who taught me the derivation for this formula, in this exact form! So Claude could simply not be more off, factually and in its characterization of the severity or what a prominent astrophysicist would say. And yes, if you're curious, I did show this exact formula in my defense, and no one had any problem with it.
So yeah I feel like that was a bit of a waste of time.
A related issue was with Genuine Technical Error #4, which complained about a related formula, measuring the duration of a supernova plateau.
It claims that my expression for the plateau duration as written incorrectly states that the plateau duration scales with the opacity of the ejecta to the one-sixth power. It then cites the original paper for this relation--which by the way, I used to write this section--and claims that my exponent is incorrect. So I trudged up to ADS, grabbed the paper, and took a look, because you know, my thesis had like fifty pages of references alone. Maybe I did make a mistake? But there it is!
I was right--Claude was just hallucinating a different exponent. The vast majority of the issues were identical to this. Just non-statements, or agreeing with me (in which case, why is it a technical error??), or just outright incorrect things.
Verdict
Overall, this was again a major disappointment. It did identify a few small typos here and there, and I appreciated one or two comments improving clarity or precision. But the overwhelming majority of the feedback was wrong in a way that made me not want to trust it in areas I'm not an expert in. It did not identify the key, major misconception I made in the supernova chapter, and it should have been super obvious to something is supposed to be "beyond PhD". It was functionally not useful to me in this way, which is honestly really surprising to me. Like, come on--this is nominally the ideal use case for LLMs. Personally, as a scientist, I would love a tool that I could run my work through and trust that it will come back with reliable, reproducible issues so I can improve and get better. And when I hear the "better than a PhD" marketing, I mean, that gets me excited! At the end of the day I just want to do the best science that we can as fast as we can so we can move into better, brighter futures. But a tool like this isn't going to speed me up or make me better if I don't trust it; if anything it slowed me down. I treated the comments as carefully as I would from a real reviewer and spent hours tracking down all thirty odd concerns it raised. And almost none of them were substantive--it was genuinely slower because it took a lot of time and extremely little came out of it. Remember, this is the easiest possible test for these claims--it is significantly easier to fact-check against legitimate sources and give feedback on writing than it is to discover something genuinely new or do an analysis from start to finish, like we've tested in our previous videos.
I also want to get out ahead of the most significant response I get when we share these results or discuss this kind of thing. I am constantly hearing that I'm not on the best model, they're getting better exponentially so who knows where they'll be in six months, that sort of thing. And I don't think this is a good line of reasoning. In these tests, I have been using the absolute highest strength model accessible on the plan that the vast majority of people will have access to. This is a fair test of the claims that are being made by the companies and their marketing teams. If I need to go up to the $200/mo Claude and use $14,000 worth of tokens just to get accurate reviews of 30 pages worth of my writing, the cost-benefit is just not there for me personally. Even at $200, on a graduate student salary, that's more than enough time for me to just talk to a professor or to just review it myself.
And I don't agree with the idea that the models are still significantly improving, at least in a way that would meaningfully change the outcome of the test we just did. They have certainly gotten significantly more capable in breadth--they have many more skills, its awesome how useful they can be when they have access to a particular tool. For example, I needed to automate some slide design thing for a powerpoint I was making and it would have been really tedious to do it by hand and also would've taken me 2 hours. But Claude did it in a few minutes. I was blown away. But that's only an "improvement" because they added the PowerPoint skill. My personal experience in terms of the model ability is that the rate of growth and improvement has slowed drastically. The performance I got out of Claude today is about the same experience I've been having with LLMs for years, and similar to the experiences of many other scientists I talk to. In fact, many journals and even the arXiv are finding that people using these tools are just overwhelming academia with quantity but the products are lacking sorely in quality, to the point where some institutions are blacklisting people who submit AI-generated papers with clear signs that they haven't been carefully checked. If AI was really beyond-PhD in reasoning and capability, would this really be the outcome?
And I'm actually trying pretty actively to find ways that these tools can speed up my workflow or help me get better as a scientist--I'm not trying to be a doomer here. They can do some things amazingly well, but for a task that requires trust and reliability and accuracy on the little things, I'm not convinced enough to rely on them yet. So just maybe next time you hear someone say that the latest AI does better than a PhD at some subject, take it with a grain of salt.
Previously on this blog:
No comments:
Post a Comment