A concept I've had to be thinking about a lot lately is something really cool called radiative levitation. This is when you have a light that's so bright that it can push atoms around and if its strong enough it can even counteract the force of gravity. You'll see this kind of concept come up not just in astrophysics but in other cool ideas like solar sails, that rely on the momentum from light to provide propulsion.
The reason I've been thinking about this is because in my research, I've been working with a system where a disk of matter is suspended on radiation coming from a compact object. I thought it would be fun to go through the math of that, and one of the easiest places to start is the Eddington luminosity, which is the maximum brightness a star can output before it starts to blow out the atoms its made of, producing something called a stellar wind. It's a really neat and straightforward derivation written down by Sir Arthur Eddington in the 1910s.
The basic idea is this. We're going to consider the simplest scenario, an electron floating around in a photon gas, bound by gravity to the star below it. This is called the classical Eddington limit. How bright do the photons have to be so that when they bump into the electron, they can transmit enough force to overcome the force of gravity? That's the fundamental question we're asking here.
Recall that force is simply a change in momentum with unit time, so we need to calculate how much momentum the photon gives to the electron. We'll assume that the post-interaction re-emitted photons don't have a preferred direction, so the average momentum of the emitted photons averages out to zero, and therefore we can assume the full incident momentum gets transferred to the electrons on average. To calculate the momentum $p$, we're going to use Einstein's equation from special relativity. You're probably familiar with the famous equation,
$$E = mc^2$$
but this is only part of the full equation, the part corresponding to the energy of a particle at rest. The full equation, where the total energy of a particle $E$ is its rest energy combined in quadrature with its energy from its momentum, is,
$$E^2 = (pc)^2 + (mc^2)^2,$$
Since a photon has no mass, $m=0$, and then $E^2 = (pc)^2 \implies E = pc$ and thus $p=E/c$. The energy of a photon comes from quantum mechanics and is simply $E_\gamma = h\nu$, where $\nu$ is the frequency of the photon. (So a higher frequency photon has more energy.) Thus our change in momentum $\delta p$ imparted to the electron from the collision is
$$\delta p \sim \frac{E_\gamma}{c}.$$
That gives us the change in momentum, but we need a timescale over which the change in momentum occurs to calculate a force, since $F = dp/dt$. When you have a gas of colliding particles—like our electron floating around in a sea of photons—one way you can think about the timescale for collisions is in terms of a mean free path $\lambda$, which is the average distance a particle will travel in a volume before it runs into another one. We can imagine that the electron will collide with a photon roughly once every mean free path, and since the photons are traveling at $c$, that collision timescale is,
$$t_{\rm coll} \sim \frac{\lambda}{c}.$$
The actual length of the path depends on two numbers, the density of the particles in the fluid and how much those particles interact. You can imagine that as the density of the particles in a volume goes up, the harder it is for particles to get through, and so the shorter the mean free path $\lambda$ is. If the number of photons in the gas per unit volume is $n_\gamma$, and the cross-sectional area that will likely produce an interaction is $\sigma_{\rm th}$ (I'll explain why its called this in a moment), then,
The reason we call this cross-sectional area $\sigma_{\rm th}$ is because it changes depending on the properties of the particles you're simulating, and this one in particular in the Thomson cross section, which is the cross-section related to electrons scattering with low-energy photons. That's why we're using it here. If the particles that are interacting have a stronger interaction than electrons and photons, for example, this area might be bigger. So now we have our collision timescale,
But remember, we're looking for a luminosity, or how bright the star has to be. So what luminosity will produce such a force? As it turns out, there's a nice way to relate luminosity and force, and it comes down to this thing called "flux". The flux $F_*$ can be thought of as measuring the amount of water flowing through a pipe with time. It's a measure of how many things pass through an area in a particular interval of time. You can think of this as sitting outside a door and counting how many people pass through in one hour. Or, another way to think about it is imagining a whole bus of people that pass through the door, bus after bus after bus. Now you don't need to count the people individually. You just need to know the density of the people on the bus, and multiply it by the velocity of the bus, which gives you the flux of people through the doorway with time. If instead of people we're talking about energetic particles in a fluid, we can use this analogy to relate the flux of energy to the energy density by
$$F_{*} = u v,$$
where $u$ is the energy density and $v$ is the velocity of the fluid. In physics terms, we would say this for a direct free streaming beam—its a little bit different if the particles are isotropically bouncing around.
So we can actually calculate this flux with all the ingredients we have above. We know the density of photons in the gas, which is $n_\gamma$. Since each photon has energy $E = h\nu$, the energy density is just $u = E n_\gamma$. And all photons move at the speed of light, so $v=c$. So we actually have this really nice relationship between flux and energy density,
$$F_* = h \nu\, n_\gamma c.$$
What can we do with this? Well, remember our force equation? If you look closely, you can see there's actually an $h\nu n_\gamma$ already in there. It turns out we were working with a flux (divided by $c$) the whole time and didn't realize it! So if substitute our new $F_*$ in, we find
Now, we can relate this directly to luminosity. A flux is the energy per unit area per unit time; a luminosity is just energy per unit time, so its just the flux added up over the whole area you're measuring (in our previous analogy, the door). In this case, its the area of the star at the radius where the electron is hovering, which we'll call $R$. So flux and luminosity are related by,
$$F_* = \frac{L}{4\pi R^2},$$
which gives us an easy way to put the luminosity in that force equation,
So we've achieved the first part of our goal, which is answering the question: how much force does a luminosity generate on an electron? To make the electron levitate, that force needs to be greater than the gravitational force of the star pulling on the electron, which we'll remember from introductory physics is just
$$F_g = \frac{GMm_e}{R^2}.$$
So, $F > F_g$ for the electron to start levitating. If we find where that limit is, we see that
And there you have it! The classical Eddington limit. Interestingly, the radius cancelled out—so it doesn't matter how big the star is, only how massive it is. If a star is brighter than this luminosity, it will be bright enough to overcome the force of gravity itself and shoot electrons off its surface! Now the final thing I want to note is a minor subtlety, which is that the Eddington limit isn't usually concerned with the electrons, but rather the protons they are attached to. But since the photons only really interact with the electrons, and the protons carry all the mass, we can just tweak this final formula to use the proton plus electron mass, and all the same logic holds.
A few days ago, I uploaded a video explaining that the Mars jellyfish anomaly was very likely to be a cosmic ray artifact due to the presence of a row of clipped pixels with no antialiasing, a sure sign of an imaging artifact--real objects, even alien spacecraft, don't produce that. I explained the mechanism for how a cosmic ray can create an artifact on an image sensor, and explicitly showed how if it impacts during a calibration frame, the subtraction can produce a black artifact on the final image. I then provided an independent confirmation of this in the form of the Mars Science Laboratory camera engineering lead explicitly describing this exact process producing a black artifact in an image from a few years ago.
Despite this, I saw a surprising number of comments that confidently and incorrectly insisted that cosmic rays can only produce WHITE artifacts, never black. Here's a few:
"As I understand it, and if the supposed experts are really experts who know what they are talking about, cosmic rays are never black."
"You may be an astrophysicist but you definitely aren't a photographer. Cosmic ray hitting a camera sensor would not leave a black spot of multiple pixels. It would excite the sensor leaving a bright white spot and it would be permanent."
"I have never seen a BLACK cosmic ray artifact. It is always WHITE."
I responded to each of these and more, reiterating the dark frame explanation from the video. But it did get me wondering--why are there so many people repeating this incorrect thing so confidently? I scratched my head about it for a while, but I didn't get a clue until a friend texted me something alarming--he had Googled a question about cosmic rays and Google gave him an answer that conflicted with my own!
I was frankly shocked by this. Cosmic rays creating black artifacts due to subtraction of calibration images is a very well-known and understood effect. It was not only confirmed by the MSL lead camera engineer but its even reproduced by amateur astronomers, like this thread on Cloudy Nights where AlphaTriplePlus says, "If you use a master dark for calibration that has a cosmic ray hit, the calibrated lights will probably exhibit a wiggly black streak as the hot pixel streak in the master dark is subtracted from every light in the live stack. Happened to me last night!"
So why did Google's AI overview get it wrong? I decided to investigate a little. I tried my friend's query and was able to reproduce what he saw. Asking "do cosmic rays cause dark artifacts?" caused Google's AI overview to confidently say, "No, cosmic rays do not cause dark artifacts", even bolding and highlighting the wrong answer! I even tried asking a few different ways--asking specifically whether cosmic rays can cause dark artifacts when impacting on a bias frame (wrong answer still!) or even going as specific as asking "if a cosmic rays strikes during calibration image acquisition, can the final image have a black artifact?" (wrong answer again!)
But after a few prompts, I did notice something interesting. Even though it always highlights and bolds the WRONG answer, sometimes, buried down deep in the answer, is the RIGHT answer, which it for some reason didn't feel was important enough to point out. For example, in the very first prompt--"can cosmic rays cause dark artifacts"--while it highlights and bolds its "NO" answer, it has a section below titled "What causes dark artifacts?" where it explicitly says that "subtracting a master dark frame can create dark "holes" or negative spots if a hot pixel was not active during the actual exposure." That's literally the explanation I gave in the original video, where the hot pixel is caused by a cosmic ray, and the correct answer to the question! But its buried down in the answer for some reason.
However, seeing this right answer pop up in the response, just buried, gave me an idea about how to get it to give the right answer. If you ask it specifically, "what artifact does a cosmic ray in a dark frame create after dark subtraction?" it will give the right answer--a negative pixel or a black hole/spot in the resulting calibrated image. So if you used AI to validate your comment by just googling and reading the first line in the Google AI overview, I hope this clears that up--its just a logical error from Google's AI overview. And I hope this makes it clear that AI makes mistakes very frequently and it is always better to try and dig up the answer yourself, or talk to an expert vs blindly accepting an AI's answer at face value, especially if the question is in a field that isn't your own. You have to be careful about interpreting responses from any AI, as exemplified by this problem we and you may have encountered trying to use AI to understand a subject as complex as image analysis of an optical camera in a very harsh planetary environment.
For the record, I tested this problem on ChatGPT and Claude as well, on the weakest models to give AI overview a fair shake. ChatGPT made a similar error but put the right answer right near the top, so I would consider that a success. If you don't read one sentence down that's kind of your fault. Claude just got the right answer out of the box. This seems to be an issue with Gemini relying too much on the fact that "cosmic rays" are most often associated with "bright artifact", but its a perfect example of how LLMs which may assume the correct answer is the one most repeated in the training data can be deeply flawed if not paired with careful reasoning capabilities and primary sources.
This also leads into something else I want to discuss, which was another theme in the comments. If AI can't be trusted--and to be clear, the answer is, it cannot--what can you trust? More importantly, I realize that I'm essentially pitting myself against Google here, which has vastly more knowledge than I have. Why should you trust me over Google Gemini, a powerful frontier AI model?
I saw this question raised MANY times in the comments. One common theme was the idea that astrophysicists don't know anything about imaging, or insinuating that astrophysics and imaging/sensing/photography are somehow disjoint mutually exclusive fields. If you think this, I can of course understand why you might be hesitant to trust my explanation when Google is very vehemently telling you the opposite. So I wanted to dive into this a bit.
First, you don't need to trust me OR Google at all. All the analysis in the previous video is analysis you can do yourselves. You can check the pixel values and see they are either (0, 0, 0) or (1, 1, 1) in the region of the artifact--that's classic clipping. You can see it goes from pitch black to sky-white with zero antialiasing, a sure sign of a sensor glitch. And if you think that I, as an astrophysicist, am not qualified to weigh in here, I showed an email from the lead engineer on the MSL camera explaining away a similar black artifact as a "textbook cosmic ray". And I assure you, the lead engineer for the MSL camera is indeed an expert on imaging.
But I would also like to speak up in defense a bit of astrophysicists everywhere. Reading comments like this, I feel like there is a severe misconception about what astrophysicists do and where our expertise lies. The vast majority of astrophysicists are observational or experimental astrophysicists, and while many of us including myself spend a lot of time working with theory and with theorists, most of our work is on data captured via IMAGING. How do you think we study the stars? We can't exactly go to them. We mostly just take pictures--sometimes fancy pictures, like what a spectrograph does, or pictures in non-optical wavelengths, like what radio telescopes do--but they are just pictures, and we analyze them and extract data from them. Even if you don't specialize in imaging, you almost certainly know fundamental principles of imaging and the details of the instrument you work on specifically, and these fundamental principles and many instrument details are shared across almost all imaging instruments. Many of us work at observatories directly and have our hands on the raw data flowing in from the telescopes. Some of us, in this case not me, but like my Astraveo colleague and good friend Dr. David James, have built the telescopes and cameras on these mountaintops with our bare hands. While I mostly work with observational data and how we can use that data to test our various theories, David literally built this enormous telescope on top of Mauna Kea in 2010! Here is David to say a few words about his thoughts on this anomaly.
I am also happy to point to my own credentials here, which I think serves as a good example of how even an astrophysicist who does not work directly on the instrument hardware must learn the details and nuances of imaging in order to do any level of data processing. In my PhD thesis alone there are countless examples of this--almost every paper I ever wrote features some discussion of handling image artifacts, and almost every paper I wrote features discussion of image artifacts specific to cosmic rays! Here are a few instances.
In any paper I wrote on imaging supernovae, we had to analyze the data with pipelines that explicitly include cosmic ray flagging and rejection. This includes recent papers from 2025 and 2026, which explicitly discuss this pipeline, and even have specific discussions of a feature in one of our spectra caused by a large, narrow cosmic-ray artifact, which is something we had to work around. In another paper, we use similar cosmic ray detection and flagging in photometry collected from many different instruments across UV optical and near-infrared wavelengths.
In one paper, we found a very early data point from ATLAS, and we weren't sure whether it was imaging noise or actually data from the supernova. It matched our model but that could have been a coincidence, and the statistical significance was only 2sigma. Using extensive analysis of the science imaging as well as control curves and the processing pipeline, we were able to moderately increase the significance to almost 3sigma but ultimately left the result inconclusive. This is a good example of the high standards we have for considering whether something is data or noise--even though we found that there was less than 1% chance that the data point was purely noise, we still determined it to be inconclusive, because with millions of data points, a less than 1 in a 100 chance isn't as unlikely as it sounds. So when you see us being cautious about artifacts or looking carefully at data, we're not just kowtowing to the "establishment", "accepting NASA's explanation", or "bending to peer pressure". We're trying to be rigorous and hold things to a high standard of discovery, because that's the only way actual science gets done.
My point with this is to say that even I, an astrophysicist who doesn't work directly on instrumentation, has had to become deeply involved and experienced with imaging algorithms and artifacts just to get my work done. So the separation that's being drawn between astrophysicists and imaging is not a very good one because that line is VERY blurry. So if you're wondering who to trust and your choices are between astrophysicists or AI--particularly the Google AI overview--I hope this gives a bit more info for you to make your decision.
I'm quite proud of this debunk--we managed to change the texture of the misinformation that was spreading about this anomaly. The things that were key for the success of this campaign were how quickly we put together a response and the very specific targeting of the communities engaging with the explanations at a friendly level. I would like to reflect more on this, and what lessons can be learned for how academia as a whole can respond to rapid misinformation/disinformation. The script of the video is below.
August 20th, 2023. Sol 3924. The Navigation Camera onboard NASA's Mars rover Curiosity photographed what appears to be an unusual structure or creature in the far distance, wandering the Martian surface. The strange shape of the object seems unlike anything seen on Mars and has recently been making the rounds on various internet forums. At first it was proposed that the object could be the NASA Ingenuity Martian helicopter, but that helicopter was over 2000 miles away when this picture was taken. UFO enthusiasts note the strange tripod like-shape, which some claim resembles a jellyfish-like being, a mechanical device, or a large black body perched atop two spindly legs; the seemingly improbable location, which has the object hovering directly above the ground in the distance; and the fact that the object is not present in an image taken earlier at the exact same location, which is being interpreted as evidence of a mobile life form or machine. What is it really?
DEBUNK
This is an older image that has been making the rounds lately, thanks to a few posts on reddit and some new amplification from tabloids like The Daily Mail, which published the responsible and level-headed headline, "Mysterious 'jellyfish' spotted in NASA Mars photo sparks theories of life on the Red Planet". Reading through the comments from UFO and alien life enthusiasts, I spotted a lot of interesting methodological problems in the way the discussions were being handled that I think are actually quite instructive. So can we figure out what this thing is, and if we can't, does that mean it is aliens?
Let's look at the evidence. First, it's objectively there. There's clearly something in the actual image from NASA that needs to be explained. This isn't an editing artifact or something someone photoshopped. Something in real life happened, and the camera recorded it. And it was something that wasn't there before, in the previous image. That's a big part I think of why this is gaining so much steam and attention.
Second, the location does seem rather odd. I saw a lot of comments saying things like, "what are the odds that a glitch or something similar would be perfectly hovering above the landscape, with its tendrils reaching down like a tripod?" And that does seem a little unusual. But without characterizing what this could be, its hard to say whether its really that unusual. And I'll to that in a minute, because I think this is an extremely common but very poorly understood source of bias in how human beings interpret low-sample data.
Finally, the shape. I'll admit, it really does look like something straight out of War of the Worlds. But this is a pretty well understood human perception thing. We like to see patterns and shapes and so we do--everywhere. Even in the random things, like the stars in the night sky, or the random inkblots of the Rorschach test.
But while the 3 things I mentioned have been shared as "evidence", they aren't really. I don't mean they don't mean aliens, I hope that's obvious at this point in the video, I mean they aren't telling you anything. The people who are commenting that its not there in the previous image and therefore it must be moving are neglecting the fact that the previous image was taken just 13 seconds earlier, which feels like if it is a big creature in the distance moving, that would be kind of fast. Especially since the "jellyfish" or "moving body" explanation would suggest its moving from right to left, pretty far away--that's a lot of ground to cover in 13 seconds, with no motion blur.
And as for the shape, a lot of what we are perceiving as the "shape" is due to intentional or accidental blurring or processing of the image as we zoom in on it or as people try to enhance it with various programs. But you can look at it pixel by pixel, using several freely available online tools. If we do this, we find that it does look weird, but it certainly doesn't look like a jellyfish or a tripod. There is an almost perfectly black line of pixels--and when I say perfectly black I mean the pixel values are (0, 0, 0) for some of these--with a fuzzy single streak looking thing below it. I don't see the "tripod" at all in this view. In fact, the two other "legs" of what I assume are the "tripod" that people are talking about are just consistent with the sky noise. Also, looking closer, unlike every other object in the image, which has some level of antialiasing (which is when the edges of real-life objects don't get perfectly terminated on a pixel, they kind of get blurred a little) the anomaly is a single row of a couple pixels perfectly terminating at the pixel boundary, without any sign of the antialiasing-like effect that a real object would have. So while there's definitely some signal here, we have been grossly misleading ourselves about the shape and structure, and even whether this is a real object in the camera.
But now we have the actual shape and structure, as we can we see here. What could cause a perfectly black row of pixels with a fuzzy streak below it? and can it explain some of the other weirdness, like the suspicious location, or the fact that its not there just a few moments earlier?
EXPLANATION
In my opinion as an astrophysicist, this is not a real object at all. I think this is just a camera glitch. And a very specific kind of camera glitch caused by something called a "cosmic ray".
See, space is filled with very high energy, fast moving particles. And sometimes--a lot of the time actually--those particles hit things. When a cosmic ray hits Earth's atmosphere, it creates a shower of secondary particles that can be detected even at the Earth's surface. We don't exactly know where they come from, but we see a LOT of them, and they seem to be coming from deep space and from the Sun.
So, how can cosmic rays create camera glitches? Well it comes down to how cameras work. Here is an overly simplified explanation. The sensor of a camera is something called a charge-coupled device, or CCD. When a photon enters a camera sensor and strikes a pixel, that photon interacts with a little piece of silicon that generates an electron. That electron is then captured by a capacitor during the exposure. The more photons interact with the little piece of silicon, the more electrons are generated and captured. When the exposure is done, the camera computer "reads out" the number of electrons counted at each pixel, and the number of electrons there are tells you the number of photons. So, more electrons, means a brighter pixel. Doing this for every pixel in the sensor across the red, green and blue pixels, gives you a whole image.
The problem is, photons aren't the only thing that can generate electrons. If a cosmic ray--which we don't care about in a picture--strikes the camera at the right angle, it can create a huge number of electrons across several pixels in the image. This is almost always far more electrons than the scene would normally generate, and so when the camera reads out the image, you'll get a big white streak that can come in many different shapes and sizes and configurations, basically anywhere in the image.
So basically what I think happened is a cosmic ray struck the camera at the exact right moment and left this big ol artifact in the image. It explains why the shape is so weird--just a perfect row of pixels with no antialiasing-- and why it wasn't present just 13 seconds earlier--a cosmic ray strike is a tiny fraction of a second event--and it also explains a few other weird things. For example, you might be wondering why is the cosmic ray artifact black here, when all the other images I've showed you so far have been white streaks. Well, NASA cameras are pretty brilliant things. As part of the process of taking an image, they take something called a "dark bias frame", which is an image taken with the camera shutter closed, to create a calibration frame that is then subtracted from the actual image. The idea here is the temperature of the environment and the camera itself can create a low-level background of photons that we don't want in our images. So by taking a picture with the shutter closed, we can measure the background that will be present in our image with the shutter open, and subtract it away to make the image more accurate.
But, if a cosmic ray strikes the camera while the bias frame is being taken, it will create a HUGE white artifact. When that artifact is then subtracted from the real image, it will leave a pitch black remnant.
The cosmic ray explanation can also explain the strange location--in that, there's nothing to explain at all. Cosmic ray hits are very frequent. They happen all the time in every instrument. They are especially frequent on Mars, which has a thin atmosphere and weaker magnetic field than Earth's, making it worse at impinging the propagation of these rays. They happen all the time, at random spots in the images, meaning that eventually, you'll get one that seems to be in a suspicious spot, purely by chance, especially if its over the sky, where a black artifact will show up more clearly. But if you look at the anomaly as an individual event and try to say "oh, what are the odds that this happened right at this perfect spot", you can fool yourself into thinking that the event is rarer than it is. Another way to think about it is like seeing a strangely shaped cloud that looks exactly like something crazy. Intuitively, we know that while it feels rare, but we also know that any given SECOND there are billions and billions of clouds each with a different shape and so suddenly a particular shape being unusual doesn't seem that unlikely.
And this something that astrophysicists have to deal with too. In a recent paper we published, we found a really strange signal--the first of its kind. We did our best to characterize it and convince ourselves that the signal was real, which wasn't so hard because the data were really high quality and the signal was very strong and occurred over a long time--over 100 days. But even though it was a very exciting find, the consensus of the community--including us, the authors--was that the only way to know for sure that this isn't a fluke is to find more of them. And even our followup work now is trying to find different ways to assess whether the signal could be real or just some type of noise. So this is something that happens all the time, and once you see how much random stuff the universe is capable of throwing at you that looks exactly like something incredible but turns out to be something kind of mundane, you sort of learn that it is much easier to trick yourself that you've proved something than it is to actually prove something beyond a shadow of a doubt. As a famous physicist once said, "The first principle is that you must not fool yourself--and that you are the easiest person to fool."
This idea of finding more similar events to show that your explanation is not a fluke leads perfectly into the final piece of this puzzle, which is that we've seen stuff like this a TON on Mars and elsewhere. Here are just a few examples that NASA themselves have published of cosmic ray streaks appearing in their cameras. Here are some that look quite similar to the strange jellyfish one we've been discussing today, complete with pitch black cores and strange dangling appendages and streaks. And to really nail it home, here is an email from Justin Maki, the engineering lead for the MSL camera, responding to a question from an enthusiast about a very similar artifact identified just a year earlier, confirming that these strange streaks are indeed cosmic rays striking the sensor during the acquisition of a dark bias frame. So that's where my money for this one, too.
A few months ago, I defended my dissertation in physics at the University of California Santa Barbara. My thesis has now been published online. You can read it here: https://escholarship.org/uc/item/6r744391 Right now, people around the world are saying AI is beyond PhD-level in knowledge and reasoning. Can it find errors in my thesis?
The cover page of my thesis, titled "Extreme Astrophysical Systems as Windows on High-Energy and Gravitational Phenomena".
In previous posts, we considered LLMs ability to make genuine scientific discoveries. It didn't really go well. Today, we're testing something way easier--is it really at a PhD level in astrophysics knowledge? I'm not going to ask it to make any discoveries--in fact, we're going to do the opposite. I'm going to give it the introductory chapters of my thesis--which contain info that should be comfortably within its training data and accessible via the internet--to see if I made any mistakes. If it finds some legitimate errors, that will be very useful to me. Will it?
LLMs are often touted for their utility in science, which they do have. But the way this is presented to the general public and as a general tool gives me some doubts. We hear all the time how publicly available LLM models can "take you to the edge of what's known in quantum physics" to make some "interesting breakthroughs" via "vibe physics" but my own experience has been that they can't make even simple leaps to discover new things. Or, we're told that they'll make science more accessible because new tools are popping up that automate the entire process of writing scientific papers from start to finish, but we consistently find that papers written using these tools are sloppy, incomplete, or don't represent meaningful advances over the literature. In fact, the only academic benchmark that LLMs seem to be having some success in is in solving graduate-level homework problems, where they are touted as scoring higher than "regular" PhDs on problem-based assessments. But safe, controlled homework problems with easily accessible solution manuals and course materials to train on is not really what makes science hard or even science.
Nonetheless, there is certainly some value in having effectively a natural-language search engine with expert-level domain knowledge in your specific subfield. I could see myself using LLMs to find papers or even to answer niche, specific questions about the physics concepts I work with and can’t find easy answers for. While, as we’ve demonstrated before, the reasoning skills of LLMs suffer greatly when pushed outside of their comfort zone slash training data, they objectively have access to vastly greater sums of knowledge than I could ever hope to achieve.
But, there is a catch. I need it to be correct. LLMs are only functionally useful to me if I can rely on them to get things right. Otherwise, I end up spending more time tracking things down and validating their output than I do just learning the thing myself. And the worst part is, I don’t feel like I even learn anything when I just double check an LLM’s work. So if it’s not correct, and it’s not saving me time, the triple whammy is I’m actually robbing myself of the chance to learn.
In our previous experiments, we gave LLMs fairly complex problems and didn’t get great results. But this time we’re asking a much simpler question—can LLMs reliably synthesize accurate information, and identify inaccuracies in my understanding? This should be much more within its wheelhouse. It’s not being asked to do anything new—just to say true things and assess whether other things are untrue. I want to say this in a different way since I think it's important to realize that the test I'm proposing here is one of the easiest tests of LLM/PhD capability claims imaginable. It is not being asked to do real science. It is not being asked to write a literature review or to even really write anything. All I want to see if it can do is fact-check in my specific domain, which by every claim and account it should be able to do (SAM ALTMAN). And even though you, the viewer, may not be an astrophysicist, I'm hoping that this kind of an assessment may help you decide whether or not to rely on info from an LLM in a domain where you don't have sufficient expertise to fact check it.
So I just finished writing my doctoral dissertation. It’s not perfect--no dissertation is. In particular, the first draft of one of my introductory chapters went through a few iterations, where one of my committee members left many thoughtful comments and suggestions for revision. It was a productive and very beneficial process and helped clear up a key misconception I had about a particular concept, as well as tidying up precision and clarity in much of the chapter.
But this gave me an idea--I have a rough draft of a chapter, and several fact-based pieces of feedback from an expert in the field. It's the kind of feedback I'd want an LLM to be able to give me, because it was about well-established concepts and knowledge that is readily accessible, not something hyper complex or requiring genuine leaps of discovery. All it requires is domain knowledge, which it should have. Remember, LLMs are better than graduate students at "everything". "In all subjects". "Simultaneously". They do well on assessments that most PhDs would fail!
Look, I don't hate this. As someone who loves learning, this would be a great tool for me to learn better and faster than I ever have before. If these things really are better than graduate level in all subjects, that includes mine and that means I basically should have a teacher that I can ping 24/7 with questions and requests for review. What kind of student wouldn't want that?
So I ran the following experiment. I gave it the rough draft of the chapter, and asked it for feedback, and then compared that feedback to the real comments an expert in the field gave me. My expectation was that it would find the most major issues that were the most obvious holes in my understanding, especially the biggest problem with the chapter. I would consider it a success if all comments it delivered were correct. I would consider it interesting if most of the comments it gave me were correct. Outright false comments--like comments that are themselves wrong or based on flawed understanding or feature hallucinations or simply don't make sense--will severely damage my trust in these tools and make it less likely that I will feel comfortable using them for work.
Let's dive into it!
The review
I used Claude Opus 4.7, the paid model, on the highest settings I could access. I have a system prompt that tries to reduce sycophancy and flattery, encouraging citations, accuracy, and truthfulness. And I didn't beat around the bush. I specifically asked it to look for problems, errors, things that are technically incorrect, that sort of thing. My hope was it would catch that key misconception I made in the chapter, the one my committee member helped me address.
Claude returned 13 of what it called "genuine technical errors", four of what it called "citation/factual issues", and five "logical/expository issues". Let's start with the technical errors. Out of all 13, one was accurate but extremely minor, three were factually correct but not an error I made--Claude simply restated something I said correctly--and 9 were fully inaccurate, hallucination-level claims. I'll give some examples of each so you can see what I'm talking about.
First, let's review the one correct thing it said.
In genuine technical error #2, Claude complained that I used the phrase "evaporated" to describe the energy being released in the form of neutrinos during the core collapse stage of a supernova. I was using it colloquially, but "evaporation" of course has an actual meaning in physics. Fair play, Claude. I did actually update my thesis to implement this correction.
But it was frankly straight downhill from here. The next genre of comments were comments that were factually correct but were listed as errors even though I couldn't figure out where the error in my statement was. Genuine Technical Error #5 claims that I did not characterize the transition between the adiabatic and non-adiabatic region of the star clearly. Basically, we think about the ejecta of a supernova in terms of how far a photon can tunnel into it from the outside, which is basically the same as asking how long does it take a photon to escape? Now, this is called "optical depth". A low optical depth means photons can escape pretty easily; a high number, means they have a harder time getting out. So here's what Claude said:
Claude's comment is this incredibly bizarre stream of consciousness. It first quotes what I said, which is that the transition between the adiabatic and non-adiabatic region occurs at optical depth of c/v or \sim30. It then agrees with me, then says "but", redoes my calculation, and then agrees with me AGAIN. It then says that's fine, but I should be explicit that this is not the photosphere, which occurs at \tau\sim1. It then acknowledges that I already mention this,
and indeed I do, right before the part it is complaining about. So what about this phrasing is slightly muddled? I think its very clear that \tau of 1 is the photosphere, and the transition between the adiabatic and non-adiabatic regions of the star is c/v\sim 30! This comment doesn't have any factual issues, but it is also saying nothing! It does this again in comment #10, where it just checks my math for the Lyapunov exponent of photon orbits around a black hole,
where it proceeds to just agree with me! Why was this listed as a genuine technical error?!
Okay let's move on to the straight up wrong stuff, which was...almost all of it. Genuine Technical Error #3 claims I got the plateau luminosity formula incorrect,
The quick overview is that some supernovae have this very long plateau of luminosity, which is created as energy deposited by the shock slowly escapes from the expanding ejecta. There's a lovely back-of-the-envelope derivation you can do to show that the luminosity of this plateau should scale approximately like this formula, L \sim \frac{E_0 c R_0}{\kappa M}.Claude is claiming this expression I used is dimensionally and physically incorrect. Let's address both of those. First, the formula makes perfect sense from a physical perspective. If you add energy to the supernova, it will get brighter, as the equation shows by having energy in the numerator. If you make the star bigger, the plateau will also be brighter. If you make the ejecta more massive, there will be more "stuff" floating around to capture photons and prevent them from escaping. We call this being more "optically thick", and when things are optically thick, photons take longer to escape. Since luminosity is energy per unit time, if you force the photons to take longer to escape without increasing the energy, the energy from the shock gets spread out over a longer timescale and the luminosity gets lower. So everything here makes sense physically.
But what about dimensionally? Claude is complaining that the units of this expression don't add up. Let's take a look. Luminosity is in units of energy per unit time, us astrophysicists like to use ergs per second. So the units on the right hand side have to multiply out to be ergs per second as well. Do they? Well, energy is in ergs. The speed of light is in centimeters per second. The radius of the star is in centimeters. The opacity is in cm^2 per gram. And finally, the mass of the star is in grams. If we multiply these all together, the units all cancel and we're left with energy per unit time. So what is Claude talking about? I have no idea.
The funniest part of this is that Claude also spotted in my thesis that Prof. Lars Bildsten is on my committee. Now, Prof. Bildsten is one of the premier theoretical physicists on the planet. His contributions to our understanding of stars are innumerable and it was a privilege to have him on my committee. Claude claims that my "mistake" here will make Lars very upset, and this error is "DANGEROUS". This is kind of hilarious, because it was in fact Prof. Lars Bildsten who taught me the derivation for this formula, in this exact form! So Claude could simply not be more off, factually and in its characterization of the severity or what a prominent astrophysicist would say. And yes, if you're curious, I did show this exact formula in my defense, and no one had any problem with it.
So yeah I feel like that was a bit of a waste of time.
A related issue was with Genuine Technical Error #4, which complained about a related formula, measuring the duration of a supernova plateau.
It claims that my expression for the plateau duration as written incorrectly states that the plateau duration scales with the opacity of the ejecta to the one-sixth power. It then cites the original paper for this relation--which by the way, I used to write this section--and claims that my exponent is incorrect. So I trudged up to ADS, grabbed the paper, and took a look, because you know, my thesis had like fifty pages of references alone. Maybe I did make a mistake? But there it is!
I was right--Claude was just hallucinating a different exponent. The vast majority of the issues were identical to this. Just non-statements, or agreeing with me (in which case, why is it a technical error??), or just outright incorrect things.
Verdict
Overall, this was again a major disappointment. It did identify a few small typos here and there, and I appreciated one or two comments improving clarity or precision. But the overwhelming majority of the feedback was wrong in a way that made me not want to trust it in areas I'm not an expert in. It did not identify the key, major misconception I made in the supernova chapter, and it should have been super obvious to something is supposed to be "beyond PhD". It was functionally not useful to me in this way, which is honestly really surprising to me. Like, come on--this is nominally the ideal use case for LLMs. Personally, as a scientist, I would love a tool that I could run my work through and trust that it will come back with reliable, reproducible issues so I can improve and get better. And when I hear the "better than a PhD" marketing, I mean, that gets me excited! At the end of the day I just want to do the best science that we can as fast as we can so we can move into better, brighter futures. But a tool like this isn't going to speed me up or make me better if I don't trust it; if anything it slowed me down. I treated the comments as carefully as I would from a real reviewer and spent hours tracking down all thirty odd concerns it raised. And almost none of them were substantive--it was genuinely slower because it took a lot of time and extremely little came out of it. Remember, this is the easiest possible test for these claims--it is significantly easier to fact-check against legitimate sources and give feedback on writing than it is to discover something genuinely new or do an analysis from start to finish, like we've tested in our previous videos.
I also want to get out ahead of the most significant response I get when we share these results or discuss this kind of thing. I am constantly hearing that I'm not on the best model, they're getting better exponentially so who knows where they'll be in six months, that sort of thing. And I don't think this is a good line of reasoning. In these tests, I have been using the absolute highest strength model accessible on the plan that the vast majority of people will have access to. This is a fair test of the claims that are being made by the companies and their marketing teams. If I need to go up to the $200/mo Claude and use $14,000 worth of tokens just to get accurate reviews of 30 pages worth of my writing, the cost-benefit is just not there for me personally. Even at $200, on a graduate student salary, that's more than enough time for me to just talk to a professor or to just review it myself.
And I don't agree with the idea that the models are still significantly improving, at least in a way that would meaningfully change the outcome of the test we just did. They have certainly gotten significantly more capable in breadth--they have many more skills, its awesome how useful they can be when they have access to a particular tool. For example, I needed to automate some slide design thing for a powerpoint I was making and it would have been really tedious to do it by hand and also would've taken me 2 hours. But Claude did it in a few minutes. I was blown away. But that's only an "improvement" because they added the PowerPoint skill. My personal experience in terms of the model ability is that the rate of growth and improvement has slowed drastically. The performance I got out of Claude today is about the same experience I've been having with LLMs for years, and similar to the experiences of many other scientists I talk to. In fact, many journals and even the arXiv are finding that people using these tools are just overwhelming academia with quantity but the products are lacking sorely in quality, to the point where some institutions are blacklisting people who submit AI-generated papers with clear signs that they haven't been carefully checked. If AI was really beyond-PhD in reasoning and capability, would this really be the outcome?
And I'm actually trying pretty actively to find ways that these tools can speed up my workflow or help me get better as a scientist--I'm not trying to be a doomer here. They can do some things amazingly well, but for a task that requires trust and reliability and accuracy on the little things, I'm not convinced enough to rely on them yet. So just maybe next time you hear someone say that the latest AI does better than a PhD at some subject, take it with a grain of salt.
As previously reported on this blog, I've been actively seeking ways to unwind and, in particular, improve my quality of sleep. I've made good progress via some of the usual suspects: blackout curtains, temperature control, limiting screen time before bed, etc. A recent very significant upgrade came when I discovered Stephen Dalton's sleep stories, which knocked me out better than any previous method I had tried. My primary method for listening to the stories was via my earbuds, but I quickly ran into problems which I described in my previous post:
They also just put me in a good mood for sleep, even if I don't use them
all the way through. More often than not these days, I don't quite fall
asleep, but I get sleepy enough, pop my earbuds back in their case
(they're a little uncomfortable to sleep with but not disruptively so)
and I'll fall straight asleep on my own. Sometimes, they slip out on
their own, and I wake up with them underneath me. I hope that won't
damage them.
I was willing to put up with the discomfort of sleeping with earbuds in because the effect of the stories was so profound. But I did wonder if maybe there was a better way. In doing some research for more comfortable (but also extremely expensive) sleep earbuds, I stumbled across this sleep mask from LC-Dolida (this post is not sponsored):
This mask is hypoallergenic, ultra-soft, and
gentle on even the most delicate skin. Sunglasses-shaped eye mask is
suitable for all face shapes, effectively blocks light, prevents light
leakage, and makes your sleep better.
Product image of the LC-Dolida sleep mask with built in bluetooth headphones.
I've previously not had great success with sleep masks, but decided to take a shot on this one, especially since the price ($32 when I purchased) was fairly non-daunting. It was well-worth the money and has been a complete game-changer.
The mask comes assembled and with a nice drawstring bag for storage, which I do use during the day. The mask is composed of two bits: a soft exoskeleton through which the wires, speakers, and control modules are arranged. The electronics can be removed, so the mask can be machine-washed, and then replaced afterwards.
The material of the mask is exceptionally light and comfortable--very plush, especially around the eyes. The raised eye cushions deform well for me as when I sleep on my side, though based on the Amazon reviews, your mileage may vary. However, I didn't have any problems using the mask while sleeping in that configuration. The hook-and-loop strap is very adjustable, but I like it somewhat tight and as a result the strap can overhang quite a bit and bunch up, which is mildly annoying at worst but not any sort of dealbreaker or anything I even notice after a few seconds.
Moving on to the electronics, I was really blown away. I was expecting the speakers to be low quality at best but they really are not. The sound is rich and well-separated with decent highs, mids, and lows. You could actively listen to music on these if you wanted. There are several buttons on the control module, accessible on the front of the mask, that control volume and play/pause/power. This is very ideal for me; one of the only downsides to listening to the sleep stories on a playlist is that they may have different volumes, and I need to handle my phone in the middle of the night, half-asleep, to correct a too-loud or too-quiet video. With the buttons built into the mask, there's no need.
The speakers are well-positioned for my ears but if they are slightly offset, you can compensate by raising the volume--they get very loud if you want them too. I usually have iOS background sounds on ("Rain" is my current go-to sound) in combination with a video from my Stephen Dalton playlist, and the mask beautifully handles both the white-noise style of the rain sounds underneath the gentle music, Stephen's voice, and whatever music and sound design he sees fit to add. It is a perfectly comfortable listening experience. The padding within the 'arms' of the mask and the small cushions on the speakers themselves make it so I cannot feel the speakers through the mask during normal use.
Maintenance is very straightforward, although it also leads me to really the only flaw of the mask. To wash the mask, simply pull the speakers and control module out of the arms and face of the mask (they are all connected) and machine-wash the mask, then replace the electronics. The speakers go in easy enough, but I cannot figure out for the life of me how to get the control module back into its place on the front of the mask. I can put it in the right spot, but it almost immediately comes out and falls to one side or the other. This makes using the buttons somewhat of a guessing game, but its not too hard. I just wish I could figure out what I did wrong--when I first got the mask, the control module was fixed in place, and I can't seem to recover that original configuration no matter how long I play around with it.
Overall, I am extremely pleased with this sleep mask. I foresee it being a permanent staple of my bedtime routine, as it solves practically all the remaining problems I was having getting to sleep. Now I get to sleep quickly and simply stay asleep, without needing to pop any earbuds out. I wake up in the morning these days finding several of Stephen's stories have been played all the way through, with absolutely no memory of any of them, which is an incredibly satisfying feeling for someone who has long struggled with sleeping deeply through the night.
Tonight, we’ll journey to the misty valleys of Ireland, where you’ll embark on a tranquil drive through the countryside in your own campervan. As rain gently taps on the roof and the mist rolls over the hills, you’ll discover the serene beauty of Ireland’s west. Feel the peaceful rhythm of the van beneath you, and allow the calming landscape to lull you into a restful sleep. 😴
My journey in attempting to relax and decompress continues. I think I've finally found a cure to my insomnia, and it's these lovely sleep stories by Stephen Dalton. Often, I have trouble sleeping, not due to any underlying medical condition (at least, I don't think so) but because my brain refuses to slow down. I feel like I spend all day red-lining in first gear and it is really hard to come down from the adrenaline and cortisol. I've been experimenting with different nighttime routines to help get my eyes off screens and allow my body to start resting, but the effect wasn't too significant. My sleep was unfortunately still light and rather patchy at best (I tracked my sleep with a smartwatch to confirm).
On a whim, I found out the meditation app I've been using (Insight Timer) had sleep stories, and I randomly decided to try one by Stephen Dalton. I was sleepily blown away. For those that don't know, these are like bedtime stories for adults. They feature soothing stories, relaxing music, tranquil sound design, and sometimes a little wind-down meditation at the beginning.
Stephen's stories seemed to work particularly well for me, I'm not quite sure why. The little relaxation session at the beginning is exactly the right tone and length to help my mind slow down a bit--like he's reaching out a hand and catching me in frantic flight, slowing me down just enough to be receptive to the story. And then the story itself just slips in and before I know it, I'm out like a light. I don't even need to track my sleep anymore--I can tell where I knocked out based on the last detail of the story I remember. I was shocked to find it was usually no more than six or seven minutes in, since it has taken me thirty to forty-five minutes to fall asleep for years. Sometimes I'm out in the relaxation session, and I don't even get to hear the story!
They also just put me in a good mood for sleep, even if I don't use them all the way through. More often than not these days, I don't quite fall asleep, but I get sleepy enough, pop my earbuds back in their case (they're a little uncomfortable to sleep with but not disruptively so) and I'll fall straight asleep on my own. Sometimes, they slip out on their own, and I wake up with them underneath me. I hope that won't damage them.
This story (A Cozy Drive Through Misty Ireland) that I posted was the first one I ever listened to, and for some reason I keep coming back to it. The piano is perfect, and the peaceful rolling thunder in the distance is so relaxing coupled with the tapping rain. In this story, you drive a camper van through the misty western mountains of Ireland. He takes you over swelling roads, by large dark loches, and along ancient stone walls marking boundaries that have stood for hundreds of years. Something about the mist and the fog and the "verdant green of the land", as Stephen describes it, fills my heart in the best way and finally gives me permission to relax.
The video is available in 4k, which seems...counterproductive.
The full list of stories I've been using is available at this YouTube playlist. I add any stories I like; though there's quite a few there, I have a few favorites, including:
The Scribe of Alexandria (absolutely love this one)
Saul the Sleepy Sloth (short but perfection)
A Magical Forest Night with a Sleepy Owl
Readings from the Shipping Forecast
Finding Harmony in the Himalayas
The Sleepy Donkey
I might write reviews for a few of these since I find the very act of reflecting on the sleep stories and writing a few words about them wonderfully relaxing in and of itself, a perfect way to wind down before sleep.